Hugging Face now hosts InternLM's AutoVerifier model, built on the Qwen3_5Moe architecture. It reads a problem, an optional reference solution, and a step‑by‑step candidate proof, then returns whether the proof is correct, any errors, and the index of the earliest mistake. The model is packaged as 40 safetensor shards totaling about 68 GiB and uses the InternS1Tokenizer. It is intended as an automatic grader for the AdvancedMathBench ProverBench suite, though the provider notes it can still make mistakes.
The model can be run without loading weights by constructing a prompt with the provided proof_verifier.md template, then feeding it to the tokenizer’s chat format.
Why it matters
It gives researchers a fast, automated way to evaluate complex math proofs without manual grading.