The ggml‑org/llama.cpp project released a new version on Oct 7, 2026. It changes the sampler so that when temperature is zero, the model always picks the highest‑probability token (greedy selection). The update keeps the existing random sampling for higher temperatures and for custom probability requests. The change applies to CPU runs and to paths that use grammar or reasoning budgets.
Why it matters
Developers can get more predictable outputs when they need deterministic text generation, without extra configuration.