Oossa
Subscribe

September 29, 2026 at 6:21 AM · 1 min read

New model predicts how AI jailbreak attacks scale with effort

Researchers led by Marco Biroli released a paper on Sep 29 2026 that offers a simple formula for estimating how often jailbreaking attempts succeed as attackers try more prompts.

Photo by FlyD on Unsplash

On Sep 29 2026, computer scientist Marco Biroli posted a paper on arXiv called “Escaping Alignment: A Physical Trap Model of Best‑of‑N Jailbreaking.” The work introduces a straightforward “barrier model” that explains how the success rate of jailbreak attacks changes when an attacker increases two variables: the number of prompt variants (N) and the number of replies per variant (M). The model uses just four numbers tied to safety mechanisms, yet it can predict outcomes for far larger attacks than have been tested so far.

What does the new model do?

The model treats each prompt as having a baseline safety level, then adds a random “barrier” that can be overcome if the attacker tries enough variations. By fitting the four numbers to data from five different language models, the authors show that all five collapse onto the same scaling curve. This means they can extrapolate from experiments with up to 100 prompt variants to scenarios with 10 000 variants, something earlier studies could not do reliably.

The paper also looks at how the generation temperature – a knob that controls randomness in AI output – affects attack success. Using the same four numbers, the authors can predict success rates at temperatures they never tested.

Why it matters for AI safety

If the model’s predictions hold up, safety engineers can estimate how many prompt variations an attacker would need before a jailbreak becomes likely. That helps them set realistic limits on how much prompting can be allowed in public APIs or how aggressively to monitor for suspicious patterns.

The work also challenges earlier claims that attack success follows a simple power‑law relationship with N. By showing a more nuanced, temperature‑dependent curve, it suggests that defenses need to consider multiple factors, not just the number of tries.

Why it matters

The model gives AI developers a quick way to gauge how hard it is for attackers to break safety guards, which can guide API limits and monitoring tools. It also warns that relying on a single metric like the number of attempts may miss other factors that make jailbreaks easier.

Was this article useful?

Sources & references

#SourceOutletDateKey takeaway
1Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking ↗arXiv cs.AISep 29, 2026arXiv:2609.32116v1 Announce Type: new Abstract: Best-of-$N$ jailbreaking (BoN) bypasses safeguards of aligned models by drawing $N$ independ

1 sources

Last updated: September 29, 2026

Oossallms.txt.md

Share

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.