The ggml‑org team updated llama.cpp (release b11518) on Oct 9 2026. The change adds 128‑bit and 96‑bit flash‑attention kernels for Metal, speeding up inference on Apple Silicon Macs. Pre‑built binaries for macOS (arm64), macOS Intel, iOS, and several Linux configurations are provided for download.
Why it matters
Apple‑silicon users can run LLMs faster without buying extra GPU hardware.