The ggml‑org team released llama.cpp version b11498 on 2026‑10‑08. The update fixes a CUDA kernel issue that occurred when certain tensor dimensions exceeded GPU grid limits. It also bundles new pre‑compiled binaries for Apple Silicon, Intel macOS, iOS, and multiple Ubuntu configurations, including CPU, Vulkan and CUDA 12.8 builds.
Why it matters
Developers can now run llama.cpp on more GPU setups without crashes and get ready‑to‑run binaries for their platform.