# Red Queen Gödel Machine lets agents and evaluators evolve together

> A new evolutionary framework lets AI agents and their learned judges improve side‑by‑side, boosting coding success from 75% to 82% on a benchmark.

Oossa · 2026-10-09 · https://oossa.com/en/red-queen-godel-machine-lets-agents-and-evaluators-evolve-together

Researchers led by Alex Iacob released the Red Queen Gödel Machine (RQGM) on arXiv. The system pairs a coding agent with a co‑evolved reviewer that grades its patches. On the DeepSWE benchmark the RQGM coder solves 82.1% of held‑out tasks, while a static‑evaluator baseline solves 75.0%. At higher effort the RQGM’s performance is close to the GPT‑6 Astra model. The framework also cuts search cost threefold for scientific writing and proof‑grading tasks.

## The facts

- Paper posted to arXiv on 2026-10-09 ("Fri Oct 09 2026").
- Coder passes 82.1% of tasks versus 75.0% baseline on DeepSWE.

## Why it matters

Developers could get AI coding assistants that keep getting better without relying on a fixed, human‑written evaluation metric.

## Sources & references

1. [The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators](https://arxiv.org/abs/2606.26294) – arXiv cs.MA, 2026-10-09

Last updated: 2026-10-09
