The ggml‑org/llama.cpp library released version b11475 on Oct 7 2026. It fixes a bug in the Metal (Apple GPU) backend where a fused MUL_MAT+ADD operation chose the wrong residual matrix, leading to incorrect results for models that use multiple matrix multiplications in one step. The fix makes the encoder pick the correct operand, restoring accurate output on Metal GPUs.
Why it matters
Developers using llama.cpp on Apple GPUs will now get reliable model outputs instead of the previously uniform probabilities.