The ggml‑org team released llama.cpp version b11507 on October 8, 2026. The update adds support for a mixture‑of‑experts (MoE) cache that can be spread over multiple GPUs. This lets developers run larger MoE models without fitting everything on a single card. The release also includes new binary packages for macOS, iOS, and several Linux configurations, including CUDA 12.8 support.
Why it matters
Developers can now run bigger, more capable AI models on multi‑GPU rigs, expanding what can be done on personal hardware.