# llama.cpp adds few‑row MMA kernel for many quantized types

> The Oct 7 2026 update expands the metal kernel to handle BF16, Q‑types and others, speeding up matrix‑multiply on Apple M3 Ultra devices.

Oossa · 2026-10-07 · https://oossa.com/en/llama-cpp-adds-few-row-mma-kernel-for-many-quantized-types

The llama.cpp project released version b11476 on Oct 7 2026. It adds a generic few‑row MMA (matrix‑multiply‑accumulate) kernel that works with BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0 and several IQ formats. The kernel kicks in at different row counts—5 rows for TQ2_0, 4 for BF16, 3 for MXFP4, Q2_0, Q2_K and IQ4_NL, and 2 for the others—where it outperforms the previous kernels on an Apple M3 Ultra. Benchmarks show the change speeds up operations from 0.23 s to 0.98 s at the threshold and improves timing across other row ranges.

## The facts

- Release b11476 published Oct 7 2026 ("Wed Oct 07 2026 16:55:09 GMT+0200").
- Few‑row MMA kernel now supports BF16, Q‑types, MXFP4, IQ types and runs faster on M3 Ultra.

## Why it matters

Developers can run quantized Llama models faster on Macs with M3 Ultra chips, reducing inference time for applications.

## Sources & references

1. [ggml-org/llama.cpp b11476](https://github.com/ggml-org/llama.cpp/releases/tag/b11476) – llama.cpp, 2026-10-07

Last updated: 2026-10-07
