# Jev model helps chess bot break 2200 bullet rating and improves label efficiency

> The probability‑only model Jev, distilled into a cheap evaluator, lifts a Stockfish bot past 2200 on Lichess and makes mixed labeling with Qwen3‑32B more effective.

Oossa · 2026-10-08 · https://oossa.com/en/jev-model-helps-chess-bot-break-2200-bullet-rating-and-improves-label-efficiency

Researchers at TypeSafe introduced Jev, a model that only returns calibrated probabilities. By embedding Jev’s judgments into Stockfish’s search and distilling them into a fast evaluator, the bot reached a 2200 bullet rating on Lichess. When a labeling budget is shared between Jev and the large language model Qwen3‑32B, the combined labels outperform using Qwen alone by about 9.6 Elo, and the benefit repeats on new openings.

## The facts

- Jev‑enhanced Stockfish surpassed a 2200 Lichess bullet rating.
- Averaging Jev and Qwen3‑32B labels beat Qwen‑only labeling by 9.6 Elo (95% interval 4.3–14.9).

## Why it matters

Chess developers can get stronger bots without expensive LLM calls, saving compute and labeling costs.

## Sources & references

1. [From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev](https://arxiv.org/abs/2610.09188) – arXiv, 2026-10-08
2. [Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev](https://arxiv.org/abs/2610.08829) – arXiv, 2026-10-08
3. [JevForest: Path Voting for Budgeted Feature Acquisition](https://arxiv.org/abs/2610.10615) – arXiv stat.ML, 2026-10-09
4. [Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged?](https://arxiv.org/abs/2610.11135) – arXiv, 2026-10-09

Last updated: 2026-10-09
