# llama.cpp bringt Flash-Attention-Kernels für Metal auf den Mac

> Die Open-Source-LLM-Laufzeitumgebung hat am 9. Oktober 2026 neue Binärdateien mit Flash-Attention-Unterstützung für 128/96 für Apple Silicon veröffentlicht.

Oossa · 2026-10-09 · https://oossa.com/de/llama-cpp-adds-flash-attention-kernels-for-metal-on-macos

Das Team von ggml-org hat llama.cpp (Release b11518) am 9. Oktober 2026 aktualisiert. Die Neuerung umfasst Flash-Attention-Kernels für Metal mit 128 Bit und 96 Bit, die die Inferenz auf Macs mit Apple Silicon beschleunigen. Vorgefertigte Binärdateien für macOS (arm64), macOS Intel, iOS und mehrere Linux-Konfigurationen stehen zum Download bereit.

## Die Fakten

- Release-Tag b11518 veröffentlicht am Fr., 9. Okt. 2026
- Ergänzt Flash-Attention-Kernels für Metal (128/96) für Apple Silicon

## Warum es wichtig ist

Nutzer von Apple-Silicon-Geräten können LLMs schneller ausführen, ohne zusätzliche GPU-Hardware kaufen zu müssen.

## Quellen und Referenzen

1. [ggml-org/llama.cpp b11518](https://github.com/ggml-org/llama.cpp/releases/tag/b11518) – llama.cpp, 2026-10-09

Zuletzt aktualisiert: 2026-10-09
