Cloudflare ran an experiment where large AI models tried to break its Web Application Firewall (WAF). The models started from payloads the firewall already blocked, then mutated them—changing encoding, placement or delivery—based on how earlier attempts were handled. The process was black‑box: the models never saw Cloudflare’s rule set or internal signals.
What the test produced
A Python harness orchestrated the test. It built HTTP requests, kept state, enforced limits and collected responses while the AI suggested the next mutation and evaluated the result. Over 45 scenarios the system generated 1,107 mutation attempts. After human triage, 49 findings were deemed worth further work, 48 of them involving command injection or server‑side request forgery (SSRF). The exercise prompted three changes to Cloudflare’s Managed Ruleset: two new detections—SSRF – Obfuscated Host and SSRF – Restricted Protocol—and an improvement to the existing SSRF – Cloud rule.
How the approach fits elsewhere
Cloudflare’s setup mirrors other security‑focused harnesses. Google’s Mandiant Agentic Vulnerability Discovery Harness chains specialized agents for code analysis, hypothesis generation and verification. OpenAI’s Codex Security and Google’s PageBreak also use a model‑in‑the‑loop but keep execution tightly constrained and require human review before any fix is applied.
Why it matters
For website owners using Cloudflare, the AI‑driven testing means the firewall now blocks a few more sneaky SSRF tricks that previously slipped through. It shows that automated, model‑guided testing can uncover subtle attack variations without exposing the underlying rule set. However, the final decision still rests with human analysts, so the speed of new protections may depend on review capacity.