Oossa

Jev model helps chess bot break 2200 bullet rating and improves label efficiency

The probability‑only model Jev, distilled into a cheap evaluator, lifts a Stockfish bot past 2200 on Lichess and makes mixed labeling with Qwen3‑32B more effective.

NoteBy Published by Oossa: Last updated: 1 min read

Researchers at TypeSafe introduced Jev, a model that only returns calibrated probabilities. By embedding Jev’s judgments into Stockfish’s search and distilling them into a fast evaluator, the bot reached a 2200 bullet rating on Lichess. When a labeling budget is shared between Jev and the large language model Qwen3‑32B, the combined labels outperform using Qwen alone by about 9.6 Elo, and the benefit repeats on new openings.

Why it matters

Chess developers can get stronger bots without expensive LLM calls, saving compute and labeling costs.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.