The ggml‑org team pushed llama.cpp version b11493 on Oct 8 2026. The update adds a SYCL implementation of grouped Mixture‑of‑Experts (MoE) XMX GEMM, a matrix multiply used in large language models. The release also bundles pre‑built binaries for macOS, iOS, Ubuntu (CPU, Vulkan and CUDA).
Why it matters
Developers can now run MoE‑based models on SYCL‑compatible GPUs, expanding hardware options for open‑source LLM workloads.