Researchers at TypeSafe introduced Jev, a model that only returns calibrated probabilities. By embedding Jev’s judgments into Stockfish’s search and distilling them into a fast evaluator, the bot reached a 2200 bullet rating on Lichess. When a labeling budget is shared between Jev and the large language model Qwen3‑32B, the combined labels outperform using Qwen alone by about 9.6 Elo, and the benefit repeats on new openings.
Why it matters
Chess developers can get stronger bots without expensive LLM calls, saving compute and labeling costs.