Oossa

llama.cpp update fixes Metal matrix‑multiply‑add fusion bug

The Oct 7 2026 release corrects a Metal backend error that caused wrong probabilities in some models.

NoteBy Published by Oossa: 1 min read

The ggml‑org/llama.cpp library released version b11475 on Oct 7 2026. It fixes a bug in the Metal (Apple GPU) backend where a fused MUL_MAT+ADD operation chose the wrong residual matrix, leading to incorrect results for models that use multiple matrix multiplications in one step. The fix makes the encoder pick the correct operand, restoring accurate output on Metal GPUs.

Why it matters

Developers using llama.cpp on Apple GPUs will now get reliable model outputs instead of the previously uniform probabilities.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.