Oossa

llama.cpp adds greedy sampling for zero‑temperature runs

The open‑source Llama.cpp library now uses greedy selection when temperature is set to zero, simplifying low‑variance generation.

NoteBy Published by Oossa: 1 min read

The ggml‑org/llama.cpp project released a new version on Oct 7, 2026. It changes the sampler so that when temperature is zero, the model always picks the highest‑probability token (greedy selection). The update keeps the existing random sampling for higher temperatures and for custom probability requests. The change applies to CPU runs and to paths that use grammar or reasoning budgets.

Why it matters

Developers can get more predictable outputs when they need deterministic text generation, without extra configuration.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.