Oossa

Red Queen Gödel Machine lets agents and evaluators evolve together

A new evolutionary framework lets AI agents and their learned judges improve side‑by‑side, boosting coding success from 75% to 82% on a benchmark.

NoteBy Published by Oossa: 1 min read

Researchers led by Alex Iacob released the Red Queen Gödel Machine (RQGM) on arXiv. The system pairs a coding agent with a co‑evolved reviewer that grades its patches. On the DeepSWE benchmark the RQGM coder solves 82.1% of held‑out tasks, while a static‑evaluator baseline solves 75.0%. At higher effort the RQGM’s performance is close to the GPT‑6 Astra model. The framework also cuts search cost threefold for scientific writing and proof‑grading tasks.

Why it matters

Developers could get AI coding assistants that keep getting better without relying on a fixed, human‑written evaluation metric.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.