Sakana AI released a paper in the Transactions on Machine Learning Research (TMLR) describing a new peer‑review system for large language models. The system uses three Claude‑based agents that work together to spot mistakes in research claims. In a benchmark of 1,164 known contradictions, the system flagged 73.43% of the core‑claim errors. The best earlier system only caught 14.81% of the same errors.
How the test was set up
The benchmark, called the Contradiction Benchmark, contains 1,164 instances where a paper’s main claim contradicts its own evidence. Sakana AI ran its Multi‑Layered Review (MLR) agents on the benchmark and recorded how many errors each system detected. The reported numbers come from Sakana AI’s own evaluation.
Why it matters
If peer‑review tools can spot most claim errors, researchers may get quicker feedback before publishing. For readers, it could mean fewer papers with hidden contradictions. Independent verification of these results is still needed.