The ggml‑org team released llama.cpp version b11391 on 4 October 2026. The update moves a CUDA variable called blocks_per_col to where it is used, a change signed off by Adrien Gallouët. The release also bundles new pre‑compiled binaries for macOS Apple Silicon, macOS Intel, iOS, and multiple Ubuntu builds (CPU, Vulkan and CUDA 12). All files are available for download from the GitHub release page.
Why it matters
Developers can now run llama.cpp on more hardware out‑of‑the‑box, especially those using CUDA‑enabled GPUs.