# llama.cpp adds flash‑attention kernels for Metal on macOS

> The open‑source LLM runtime released new binaries with 128/96 flash‑attention support for Apple Silicon on Oct 9 2026.

Oossa · 2026-10-09 · https://oossa.com/en/llama-cpp-adds-flash-attention-kernels-for-metal-on-macos

The ggml‑org team updated llama.cpp (release b11518) on Oct 9 2026. The change adds 128‑bit and 96‑bit flash‑attention kernels for Metal, speeding up inference on Apple Silicon Macs. Pre‑built binaries for macOS (arm64), macOS Intel, iOS, and several Linux configurations are provided for download.

## The facts

- Release tag b11518 published on Fri Oct 09 2026
- Adds Metal flash‑attention kernels (128/96) for Apple Silicon

## Why it matters

Apple‑silicon users can run LLMs faster without buying extra GPU hardware.

## Sources & references

1. [ggml-org/llama.cpp b11518](https://github.com/ggml-org/llama.cpp/releases/tag/b11518) – llama.cpp, 2026-10-09

Last updated: 2026-10-09
