Oossa

llama.cpp v0.1.11534 adds CUDA copy optimizations

The open‑source LLaMA inference library removes redundant GPU memory moves, cutting overhead for certain workloads.

NoteBy Published by Oossa: 1 min read

The ggml‑org team released llama.cpp version b11534 on 2026-10-09. The update fuses state‑snapshot copies into the recurrent cache and drops extra CUDA copies when K = 1, a non‑speculative decoding case. The change is limited to the CUDA backend, so CPU and Vulkan builds are unchanged. Pre‑built binaries for macOS, iOS, and several Ubuntu flavors are available for download.

Users running the CUDA version should see slightly lower GPU memory traffic, which can improve speed on supported hardware.

Why it matters

CUDA users may get faster inference because the library now moves less data on the GPU.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.