The ggml‑org team published llama.cpp version b11540 on Oct 10 2026. The update adds SYCL support to speed up MXFP4 mixture‑of‑experts (MoE) models using arithmetic decoding and weight reordering. It also ships pre‑built binaries for macOS Apple Silicon, Intel, iOS, and several Ubuntu variants (CPU, Vulkan and CUDA). The changes are listed under the tag b11540 on GitHub.
Why it matters
Developers can now run larger MoE models faster on GPUs that support SYCL, expanding hardware options for llama.cpp users.