The ggml‑org/llama.cpp project released version b11530 on 2026‑10‑09. The change keeps the backend sampling graph static across ubatches, so the graph no longer changes shape during reserve and decode steps. This fixes crashes that happened when the graph topology shifted after a reserve build. The fix applies to all listed builds – macOS, Linux, Android, Windows and openEuler – and to all hardware backends such as CUDA, Vulkan and OpenCL.
Why it matters
Developers using llama.cpp will see fewer runtime errors when running multi‑output generations on any supported CPU or GPU.