# llama.cpp adds greedy sampling for zero‑temperature runs

> The open‑source Llama.cpp library now uses greedy selection when temperature is set to zero, simplifying low‑variance generation.

Oossa · 2026-10-07 · https://oossa.com/en/llama-cpp-adds-greedy-sampling-for-zero-temperature-runs

The ggml‑org/llama.cpp project released a new version on Oct 7, 2026. It changes the sampler so that when temperature is zero, the model always picks the highest‑probability token (greedy selection). The update keeps the existing random sampling for higher temperatures and for custom probability requests. The change applies to CPU runs and to paths that use grammar or reasoning budgets.

## The facts

- Release date: Oct 7, 2026
- New feature: greedy selection for temperature‑zero chains

## Why it matters

Developers can get more predictable outputs when they need deterministic text generation, without extra configuration.

## Sources & references

1. [ggml-org/llama.cpp b11472](https://github.com/ggml-org/llama.cpp/releases/tag/b11472) – llama.cpp, 2026-10-07

Last updated: 2026-10-07
