# Aleph Alpha trains LLM to flag unsupported answers

> New Merlin‑Arthur training lets a model say “I don’t know” when a document doesn’t contain the answer, cutting hallucinations by up to 35 points.

Oossa · 2026-09-29 · https://oossa.com/en/aleph-alpha-trains-llm-to-flag-unsupported-answers

Aleph Alpha released a method that teaches a language model to refuse answering when the supplied document lacks the needed information. The technique, called a Merlin‑Arthur protocol, pits the model (Arthur) against two helpers: Merlin, who supplies a complete passage, and Morgana, who removes the decisive sentence. During training Arthur never knows which side he faces, so he learns to answer only when the evidence is present and to abstain otherwise. Tests on five QA benchmarks show wrong answers drop by up to 35 percentage points and a new “grounding score” rises by as much as 0.38.

## The facts

- Paper posted on arXiv as 2512.11614 (August 8 2026)
- Wrong answers fell up to 35 pp; grounding score improved up to 0.38

## Why it matters

It gives users a reliable way to know when an AI’s answer truly comes from their own documents, reducing risky hallucinations.

## Sources & references

1. [Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models](https://www.aleph-alpha.com/en/blog/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models/) – Aleph Alpha, 2026-09-29

Last updated: 2026-09-29
