The llama.cpp project released version b11476 on Oct 7 2026. It adds a generic few‑row MMA (matrix‑multiply‑accumulate) kernel that works with BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0 and several IQ formats. The kernel kicks in at different row counts—5 rows for TQ2_0, 4 for BF16, 3 for MXFP4, Q2_0, Q2_K and IQ4_NL, and 2 for the others—where it outperforms the previous kernels on an Apple M3 Ultra. Benchmarks show the change speeds up operations from 0.23 s to 0.98 s at the threshold and improves timing across other row ranges.
Why it matters
Developers can run quantized Llama models faster on Macs with M3 Ultra chips, reducing inference time for applications.