# llama.cpp update fixes Metal matrix‑multiply‑add fusion bug

> The Oct 7 2026 release corrects a Metal backend error that caused wrong probabilities in some models.

Oossa · 2026-10-07 · https://oossa.com/en/llama-cpp-update-fixes-metal-matrix-multiply-add-fusion-bug

The ggml‑org/llama.cpp library released version b11475 on Oct 7 2026. It fixes a bug in the Metal (Apple GPU) backend where a fused MUL_MAT+ADD operation chose the wrong residual matrix, leading to incorrect results for models that use multiple matrix multiplications in one step. The fix makes the encoder pick the correct operand, restoring accurate output on Metal GPUs.

## The facts

- Release version b11475 published Oct 7 2026
- Before the fix, 27 of 28 test cases on Metal failed

## Why it matters

Developers using llama.cpp on Apple GPUs will now get reliable model outputs instead of the previously uniform probabilities.

## Sources & references

1. [ggml-org/llama.cpp b11475](https://github.com/ggml-org/llama.cpp/releases/tag/b11475) – llama.cpp, 2026-10-07

Last updated: 2026-10-07
