Oossa

InternLM releases AutoVerifier model for grading math proofs on Hugging Face

The new model checks natural‑language proofs, points out errors and marks the first wrong step, helping evaluate AdvancedMathBench submissions.

NoteOossa1 min read

Hugging Face now hosts InternLM's AutoVerifier model, built on the Qwen3_5Moe architecture. It reads a problem, an optional reference solution, and a step‑by‑step candidate proof, then returns whether the proof is correct, any errors, and the index of the earliest mistake. The model is packaged as 40 safetensor shards totaling about 68 GiB and uses the InternS1Tokenizer. It is intended as an automatic grader for the AdvancedMathBench ProverBench suite, though the provider notes it can still make mistakes.

The model can be run without loading weights by constructing a prompt with the provided proof_verifier.md template, then feeding it to the tokenizer’s chat format.

Why it matters

It gives researchers a fast, automated way to evaluate complex math proofs without manual grading.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Sources & references

#SourceOutletDateKey takeaway
1New model on Hugging Face: internlm/AdvancedMathBench-AutoVerifier (image-text-to-text) ↗InternLMSep 29, 2026Organisation internlm published internlm/AdvancedMathBench-AutoVerifier.

1 sources

Last updated: ·Markdown·llms.txt