The ggml‑org team released llama.cpp version b11461 on 2026-10-07. The update fixes a slow checkpoint read when using Vulkan on AMD integrated GPUs. It also provides new binary packages for Apple Silicon, Intel Macs, iOS, and several Linux builds (CPU, Vulkan, CUDA 12.8). Users can download the appropriate archive from the GitHub release page.
Why it matters
AMD laptop users can now run LLaMA models faster without the previous slowdown.