ggml‑org released llama.cpp version b11462 on 2026‑10‑07. The update cleans up SYCL‑based attention buffers, a low‑level graphics‑compute interface. It also ships pre‑built binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu flavors, including CPU, Vulkan and CUDA 12.8 builds. The binaries are available for download from the GitHub release page.
Why it matters
Developers can now run llama.cpp locally on more platforms without building from source, speeding up experimentation.