The llama.cpp project released version b11430 on 2026‑10‑05. It adds a flag called GGML_HEXAGON_FA_HEAD_SPLIT that lets each core handle only a subset of attention heads when the number of KV heads divides evenly across cores. This reduces the amount of KV cache each core must read, improving memory bandwidth use. In tests on four cores, Qwen3‑0.6B saw a 58% speed boost, and llama‑3.2‑3B a 49% boost. The change is optional and controlled by the new –fa-head-split option in run.py.
Why it matters
Developers running llama.cpp on Qualcomm Hexagon hardware can get noticeably faster inference on supported models without changing their code.