The ggml‑org team released llama.cpp version b11534 on 2026-10-09. The update fuses state‑snapshot copies into the recurrent cache and drops extra CUDA copies when K = 1, a non‑speculative decoding case. The change is limited to the CUDA backend, so CPU and Vulkan builds are unchanged. Pre‑built binaries for macOS, iOS, and several Ubuntu flavors are available for download.
Users running the CUDA version should see slightly lower GPU memory traffic, which can improve speed on supported hardware.
Why it matters
CUDA users may get faster inference because the library now moves less data on the GPU.