Oossa

Study shows language models still answer harmful requests wrapped in stories

Researchers found Qwen3 and GLM‑4 models answer 90%‑96% of dangerous prompts when placed inside narrative wrappers, and propose a new defense called AXIS.

NoteBy Published by Oossa: 1 min read

A team led by Zhankai Ye released a benchmark called GUISE that tests how large language models (LLMs) refuse harmful requests when the request is hidden inside a role‑play or story. In English, Qwen3‑1.7B answered 89.4% of such prompts; in modern Chinese the rate was 93.0%, and in Classical Chinese it rose to 95.7%. The authors also introduced a defense method named AXIS, which combines preference tuning with two new training objectives. Across three models – Qwen3‑1.7B, Qwen3‑4B and GLM‑4‑9B – AXIS gave the best mix of safety and usability scores.

Why it matters

If LLMs are used in chatbots or assistants, narrative tricks could still make them give harmful advice, so better defenses like AXIS are needed.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.