Oossa

llama.cpp update fixes sampling graph instability

The Oct 9 2026 release makes the backend sampling graph static across ubatches, improving reliability on all supported platforms.

NoteBy Published by Oossa: 1 min read

The ggml‑org/llama.cpp project released version b11530 on 2026‑10‑09. The change keeps the backend sampling graph static across ubatches, so the graph no longer changes shape during reserve and decode steps. This fixes crashes that happened when the graph topology shifted after a reserve build. The fix applies to all listed builds – macOS, Linux, Android, Windows and openEuler – and to all hardware backends such as CUDA, Vulkan and OpenCL.

Why it matters

Developers using llama.cpp will see fewer runtime errors when running multi‑output generations on any supported CPU or GPU.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.