The open‑source llama.cpp library released version b11511 on Oct 8, 2026. It fixes a CUDA out‑of‑bounds read bug in the MMQ implementation. The update also adds ready‑to‑run binaries for macOS (Apple Silicon and Intel), iOS, and Ubuntu on x64, arm64 and s390x, plus Vulkan and CUDA‑12.8 builds.
Why it matters
Developers can now run llama.cpp on more hardware without crashes, especially those using NVIDIA GPUs.