Oossa

llama.cpp adds few‑row MMA kernel for many quantized types

The Oct 7 2026 update expands the metal kernel to handle BF16, Q‑types and others, speeding up matrix‑multiply on Apple M3 Ultra devices.

NoteBy Published by Oossa: 1 min read

The llama.cpp project released version b11476 on Oct 7 2026. It adds a generic few‑row MMA (matrix‑multiply‑accumulate) kernel that works with BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0 and several IQ formats. The kernel kicks in at different row counts—5 rows for TQ2_0, 4 for BF16, 3 for MXFP4, Q2_0, Q2_K and IQ4_NL, and 2 for the others—where it outperforms the previous kernels on an Apple M3 Ultra. Benchmarks show the change speeds up operations from 0.23 s to 0.98 s at the threshold and improves timing across other row ranges.

Why it matters

Developers can run quantized Llama models faster on Macs with M3 Ultra chips, reducing inference time for applications.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.