# llama.cpp adds head‑split option for faster flash‑attention on Hexagon

> The new GGML_HEXAGON_FA_HEAD_SPLIT flag speeds up multi‑core flash‑attention by up to 58% on supported models.

Oossa · 2026-10-05 · https://oossa.com/en/llama-cpp-adds-head-split-option-for-faster-flash-attention-on-hexagon

The llama.cpp project released version b11430 on 2026‑10‑05. It adds a flag called GGML_HEXAGON_FA_HEAD_SPLIT that lets each core handle only a subset of attention heads when the number of KV heads divides evenly across cores. This reduces the amount of KV cache each core must read, improving memory bandwidth use. In tests on four cores, Qwen3‑0.6B saw a 58% speed boost, and llama‑3.2‑3B a 49% boost. The change is optional and controlled by the new –fa-head-split option in run.py.

## The facts

- Release date: 2026‑10‑05 ("Mon Oct 05 2026 22:38:39 GMT+0200")
- Speedup: Qwen3‑0.6B 6977 → 11026 tokens/s (+58%)

## Why it matters

Developers running llama.cpp on Qualcomm Hexagon hardware can get noticeably faster inference on supported models without changing their code.

## Sources & references

1. [ggml-org/llama.cpp b11430](https://github.com/ggml-org/llama.cpp/releases/tag/b11430) – llama.cpp, 2026-10-05

Last updated: 2026-10-05
