Researchers led by Alex Iacob released the Red Queen Gödel Machine (RQGM) on arXiv. The system pairs a coding agent with a co‑evolved reviewer that grades its patches. On the DeepSWE benchmark the RQGM coder solves 82.1% of held‑out tasks, while a static‑evaluator baseline solves 75.0%. At higher effort the RQGM’s performance is close to the GPT‑6 Astra model. The framework also cuts search cost threefold for scientific writing and proof‑grading tasks.
Why it matters
Developers could get AI coding assistants that keep getting better without relying on a fixed, human‑written evaluation metric.