The ggml‑org team posted llama.cpp release b11515 on 2026‑10‑09. The update adds a SYCL‑based Q5_K reorder‑layout matrix‑multiply‑vector‑quantization (MMVQ) and a fused GLU operator, improving performance on compatible hardware. The release also bundles ready‑to‑run binaries for Apple Silicon, Intel Macs, iOS, Ubuntu (CPU, Vulkan, CUDA 12.8) and IBM s390x. Users can download the appropriate archive from the GitHub page.
Why it matters
Developers can now run llama.cpp faster on SYCL‑compatible GPUs without building from source.