The ggml‑org team released llama.cpp version b11480 on 2026‑10‑08. The update adds a GPU cache for MoE experts that lives in host memory, helping large models run faster on GPUs. The change is marked as assisted‑by Claude and uses a new llama_moe_cache_ptr API. Pre‑built binaries for macOS, iOS, and several Ubuntu configurations are available for download.
Why it matters
Developers can now run MoE models more efficiently on consumer GPUs, lowering the hardware cost for AI projects.