Researchers have discovered that large language models store a geometric cue for whether an answer is right. By looking at the average shift in hidden states from wrong to right answers, they can create a scoring vector that ranks candidates without any fine‑tuning. Tested on five models from 1 billion to 8 billion parameters, the technique beats the usual zero‑shot log‑probability approach by up to 32 points on ARC‑Challenge and MMLU and by up to 52 points on TruthfulQA.
How the method works
The authors take fifty labeled examples, run them through the model to collect hidden representations about halfway down the network, and compute the mean vector that points from incorrect to correct answer states. At inference time each answer candidate only needs one forward pass and a single dot product with this vector. No answer text is generated, and no model parameters are changed.
Implications for model reliability
When used as a hallucination detector, the geometric direction reaches an AUROC of 0.693, compared with 0.578 for standard probability scoring. The study also shows that correctness cues for factual reasoning, domain knowledge, and truthfulness occupy nearly orthogonal subspaces, meaning models keep separate signals for different kinds of correctness. This may explain why models often know the right answer internally but fail to output it.
Why it matters
For everyday users, the technique could lead to more trustworthy answers from chatbots, especially on factual questions. It also suggests a way to flag likely hallucinations without expensive computation. However, the approach still requires access to hidden model states, which most public APIs do not expose.