# llama.cpp update improves MoE matrix multiplication efficiency

> The Sep 30, 2026 release tweaks Vulkan tile selection, cutting wasted work in Mixture‑of‑Experts models.

Oossa · 2026-09-29 · https://oossa.com/en/llama-cpp-update-improves-moe-matrix-multiplication-efficiency

The open‑source llama.cpp library added a Vulkan fix for MoE‑aware matrix multiplication. The change adjusts how the code picks tile sizes for the mat_mul_id operation, matching the true number of active rows per expert. In a test on a 30‑billion‑parameter model, the fix reduced idle GPU work that had taken 55% of the run time.

## The facts

- Release version b11265 posted Sep 30, 2026
- Performance waste dropped from 55% to near zero in the Sarvam 30B test

## Why it matters

Less idle GPU time means faster inference for large MoE models running on Vulkan‑compatible hardware.

## Sources & references

1. [ggml-org/llama.cpp b11265](https://github.com/ggml-org/llama.cpp/releases/tag/b11265) – llama.cpp, 2026-09-29

Last updated: 2026-09-29
