# llama.cpp adds GPU cache for MoE experts

> The open‑source llama.cpp library now stores mixture‑of‑experts (MoE) data in a GPU cache kept in host memory.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-adds-gpu-cache-for-moe-experts

The ggml‑org team released llama.cpp version b11480 on 2026‑10‑08. The update adds a GPU cache for MoE experts that lives in host memory, helping large models run faster on GPUs. The change is marked as assisted‑by Claude and uses a new llama_moe_cache_ptr API. Pre‑built binaries for macOS, iOS, and several Ubuntu configurations are available for download.

## The facts

- Release version: b11480
- Release date: 2026‑10‑08

## Why it matters

Developers can now run MoE models more efficiently on consumer GPUs, lowering the hardware cost for AI projects.

## Sources & references

1. [ggml-org/llama.cpp b11480](https://github.com/ggml-org/llama.cpp/releases/tag/b11480) – llama.cpp, 2026-10-08
2. [Cascadia: Resident 975B MoE Inference on Eleven AI PCs](https://arxiv.org/abs/2610.07219) – arXiv, 2026-10-07

Last updated: 2026-10-08
