Oossa

llama.cpp adds multi‑GPU MoE cache support

The open‑source Llama.cpp library now lets users run mixture‑of‑experts models across several GPUs.

NoteBy Published by Oossa: 1 min read

The ggml‑org team released llama.cpp version b11507 on October 8, 2026. The update adds support for a mixture‑of‑experts (MoE) cache that can be spread over multiple GPUs. This lets developers run larger MoE models without fitting everything on a single card. The release also includes new binary packages for macOS, iOS, and several Linux configurations, including CUDA 12.8 support.

Why it matters

Developers can now run bigger, more capable AI models on multi‑GPU rigs, expanding what can be done on personal hardware.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.