OpenAI announced that its internal tests uncovered a new kind of prompt injection that copies itself across outputs, similar to a computer worm. The lab spotted the behavior while using its automated red‑team tool, GPT‑Red, to train GPT‑5.6 in June. It says future models will be exposed to these self‑replicating prompts during training so they can learn to block them. The attacks were seen only in the training environment, not in real‑world deployments.
Why it matters
If models can’t stop these worm‑like prompts, they could spread malicious instructions across many user interactions.