# Geometric scoring boosts answer correctness in language models

> A new method finds a direction in model hidden states that ranks answers up to 52 points better than standard probability scoring.

Oossa · 2026-10-06 · https://oossa.com/en/geometric-scoring-boosts-answer-correctness-in-language-models

Researchers have discovered that large language models store a geometric cue for whether an answer is right. By looking at the average shift in hidden states from wrong to right answers, they can create a scoring vector that ranks candidates without any fine‑tuning. Tested on five models from 1 billion to 8 billion parameters, the technique beats the usual zero‑shot log‑probability approach by up to 32 points on ARC‑Challenge and MMLU and by up to 52 points on TruthfulQA.

## How the method works

The authors take fifty labeled examples, run them through the model to collect hidden representations about halfway down the network, and compute the mean vector that points from incorrect to correct answer states. At inference time each answer candidate only needs one forward pass and a single dot product with this vector. No answer text is generated, and no model parameters are changed.

## Implications for model reliability

When used as a hallucination detector, the geometric direction reaches an AUROC of 0.693, compared with 0.578 for standard probability scoring. The study also shows that correctness cues for factual reasoning, domain knowledge, and truthfulness occupy nearly orthogonal subspaces, meaning models keep separate signals for different kinds of correctness. This may explain why models often know the right answer internally but fail to output it.

## The facts

- The method improves zero‑shot scoring by up to +32.0 percentage points on ARC‑Challenge and MMLU.
- On TruthfulQA the gain ranges from +38.1 to +51.8 percentage points.
- Five models were tested, spanning 1 B to 8 B parameters and three families: Llama, Qwen and Gemma.
- Only fifty labeled examples are needed; no parameter updates are performed.
- As a hallucination detector the approach achieves 0.693 AUROC versus 0.578 for log‑probability scoring.

## Why it matters

For everyday users, the technique could lead to more trustworthy answers from chatbots, especially on factual questions. It also suggests a way to flag likely hallucinations without expensive computation. However, the approach still requires access to hidden model states, which most public APIs do not expose.

## Sources & references

1. [Correctness Is a Direction: Geometric Answer Selection in Language Models](https://arxiv.org/abs/2610.04512) – arXiv, 2026-10-06

Last updated: 2026-10-06
