OossaAI is evolving fast. We explain it simply.

OpenAI reports self‑replicating prompt injection tests in its GPT models

OpenAI says it found worm‑like prompt attacks in internal testing and is training future models with its red‑team bot to recognize them.

NoteOossa1 min read

OpenAI announced that its internal tests uncovered a new kind of prompt injection that copies itself across outputs, similar to a computer worm. The lab spotted the behavior while using its automated red‑team tool, GPT‑Red, to train GPT‑5.6 in June. It says future models will be exposed to these self‑replicating prompts during training so they can learn to block them. The attacks were seen only in the training environment, not in real‑world deployments.

Why it matters

If models can’t stop these worm‑like prompts, they could spread malicious instructions across many user interactions.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Sources & references

#SourceOutletDateKey takeaway
1Add one more AI worry to the nightmare scenario: self-replicating prompt injections ↗The RegisterSep 29, 2026It's a worm attack, AI-style

1 sources

Last updated: ·Markdown·llms.txt