# llama.cpp adds multi‑GPU MoE cache support

> The open‑source Llama.cpp library now lets users run mixture‑of‑experts models across several GPUs.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-adds-multi-gpu-moe-cache-support

The ggml‑org team released llama.cpp version b11507 on October 8, 2026. The update adds support for a mixture‑of‑experts (MoE) cache that can be spread over multiple GPUs. This lets developers run larger MoE models without fitting everything on a single card. The release also includes new binary packages for macOS, iOS, and several Linux configurations, including CUDA 12.8 support.

## The facts

- Release version b11507 published on 2026-10-08
- Adds MoE cache across multiple GPUs

## Why it matters

Developers can now run bigger, more capable AI models on multi‑GPU rigs, expanding what can be done on personal hardware.

## Sources & references

1. [ggml-org/llama.cpp b11507](https://github.com/ggml-org/llama.cpp/releases/tag/b11507) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
