Oossa

llama.cpp adds flash‑attention kernels for Metal on macOS

The open‑source LLM runtime released new binaries with 128/96 flash‑attention support for Apple Silicon on Oct 9 2026.

NoteBy Published by Oossa: 1 min read

The ggml‑org team updated llama.cpp (release b11518) on Oct 9 2026. The change adds 128‑bit and 96‑bit flash‑attention kernels for Metal, speeding up inference on Apple Silicon Macs. Pre‑built binaries for macOS (arm64), macOS Intel, iOS, and several Linux configurations are provided for download.

Why it matters

Apple‑silicon users can run LLMs faster without buying extra GPU hardware.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.