# llama.cpp v0.2.0 pre‑release adds Hexagon CPU optimizations

> The open‑source LLM runner updates its Hexagon backend for faster concat operations on Snapdragon chips.

Oossa · 2026-09-30 · https://oossa.com/en/llama-cpp-v0-2-0-pre-release-adds-hexagon-cpu-optimizations

The ggml‑org/llama.cpp project published a pre‑release build (tag b11272) on 30 Sept 2026. It adds a Hexagon‑specific optimization for the concat operation, which is a common step in language‑model inference. The change reduces packet overhead and uses a faster division routine, speeding up the hot loop that moves data around. The update is part of a broader set of builds for many platforms, including Android Snapdragon devices with Hexagon NPUs.

## The facts

- Release tag: b11272 (commit 7fee178) published 30 Sep 2026
- Hexagon concat optimization cuts packet count and uses fast‑div instruction

## Why it matters

It makes running LLMs on Snapdragon phones noticeably faster, lowering latency for on‑device AI tasks.

## Sources & references

1. [ggml-org/llama.cpp b11272](https://github.com/ggml-org/llama.cpp/releases/tag/b11272) – llama.cpp, 2026-09-30

Last updated: 2026-09-30
