OossaAI is evolving fast. We explain it simply.

llama.cpp update improves MoE matrix multiplication efficiency

The Sep 30, 2026 release tweaks Vulkan tile selection, cutting wasted work in Mixture‑of‑Experts models.

NoteOossa1 min read

The open‑source llama.cpp library added a Vulkan fix for MoE‑aware matrix multiplication. The change adjusts how the code picks tile sizes for the mat_mul_id operation, matching the true number of active rows per expert. In a test on a 30‑billion‑parameter model, the fix reduced idle GPU work that had taken 55% of the run time.

Why it matters

Less idle GPU time means faster inference for large MoE models running on Vulkan‑compatible hardware.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Sources & references

#SourceOutletDateKey takeaway
1ggml-org/llama.cpp b11265 ↗llama.cppSep 29, 2026<details open> vulkan: MOE aware mat_mul_id tile selection (#29182) mut_mul_id selected its matmul tile with total token count.

1 sources

Last updated: ·Markdown·llms.txt