Oossa

Aleph Alpha trains LLM to flag unsupported answers

New Merlin‑Arthur training lets a model say “I don’t know” when a document doesn’t contain the answer, cutting hallucinations by up to 35 points.

NoteOossa1 min read

Aleph Alpha released a method that teaches a language model to refuse answering when the supplied document lacks the needed information. The technique, called a Merlin‑Arthur protocol, pits the model (Arthur) against two helpers: Merlin, who supplies a complete passage, and Morgana, who removes the decisive sentence. During training Arthur never knows which side he faces, so he learns to answer only when the evidence is present and to abstain otherwise. Tests on five QA benchmarks show wrong answers drop by up to 35 percentage points and a new “grounding score” rises by as much as 0.38.

Why it matters

It gives users a reliable way to know when an AI’s answer truly comes from their own documents, reducing risky hallucinations.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Sources & references

1 sources

Last updated: ·Markdown·llms.txt