The open‑source llama.cpp library added a Vulkan fix for MoE‑aware matrix multiplication. The change adjusts how the code picks tile sizes for the mat_mul_id operation, matching the true number of active rows per expert. In a test on a 30‑billion‑parameter model, the fix reduced idle GPU work that had taken 55% of the run time.
Why it matters
Less idle GPU time means faster inference for large MoE models running on Vulkan‑compatible hardware.