Oossa

llama.cpp adds head‑split option for faster flash‑attention on Hexagon

The new GGML_HEXAGON_FA_HEAD_SPLIT flag speeds up multi‑core flash‑attention by up to 58% on supported models.

NoteOossaPublished by Oossa: 1 min read

The llama.cpp project released version b11430 on 2026‑10‑05. It adds a flag called GGML_HEXAGON_FA_HEAD_SPLIT that lets each core handle only a subset of attention heads when the number of KV heads divides evenly across cores. This reduces the amount of KV cache each core must read, improving memory bandwidth use. In tests on four cores, Qwen3‑0.6B saw a 58% speed boost, and llama‑3.2‑3B a 49% boost. The change is optional and controlled by the new –fa-head-split option in run.py.

Why it matters

Developers running llama.cpp on Qualcomm Hexagon hardware can get noticeably faster inference on supported models without changing their code.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.