# Sakana AI’s review AI catches 73% of claim errors in test

> The company’s new three‑agent Claude‑based reviewer found 73.43% of core‑claim errors, far above the previous best of 14.81%.

Oossa · 2026-10-11 · https://oossa.com/en/sakana-ai-s-review-ai-catches-73-of-claim-errors-in-test

Sakana AI released a paper in the Transactions on Machine Learning Research (TMLR) describing a new peer‑review system for large language models. The system uses three Claude‑based agents that work together to spot mistakes in research claims. In a benchmark of 1,164 known contradictions, the system flagged 73.43% of the core‑claim errors. The best earlier system only caught 14.81% of the same errors.

## How the test was set up

The benchmark, called the Contradiction Benchmark, contains 1,164 instances where a paper’s main claim contradicts its own evidence. Sakana AI ran its Multi‑Layered Review (MLR) agents on the benchmark and recorded how many errors each system detected. The reported numbers come from Sakana AI’s own evaluation.

## The facts

- Sakana AI’s Multi‑Layered Review uses three Claude‑based reviewer agents.
- The system was tested on a 1,164‑error Contradiction Benchmark.
- MLR caught 73.43% of core‑claim errors.
- The prior best system caught only 14.81% of those errors.
- The results were published in a TMLR paper on Oct 10 2026.

## Why it matters

If peer‑review tools can spot most claim errors, researchers may get quicker feedback before publishing. For readers, it could mean fewer papers with hidden contradictions. Independent verification of these results is still needed.

## Sources & references

1. [Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors](https://www.marktechpost.com/2026/10/10/sakana-ais-llm-peer-review-system-catches-73-of-core-claim-errors/) – MarkTechPost, 2026-10-10

Last updated: 2026-10-11
