Oossa

llama.cpp adds Hexagon buffer support and larger default allocation

The open‑source llama.cpp library now lets Hexagon devices split large tensors and raises the default dynamic buffer to 512 MB.

NoteBy Published by Oossa: 1 min read

The llama.cpp project released version b11490 on 2026‑10‑08. It adds support for the Hexagon "alloc_buffer_n" API, allowing large tensors to be split across multiple buffers. The default dynamic buffer size is increased from 128 MB to 512 MB to avoid performance drops with big models. A new "--no-embd-offload" flag simplifies command lines for devices that cannot offload embeddings. The changes were co‑authored by Jhen‑Jie Hong.

Why it matters

Developers using Hexagon‑based hardware can run larger language models more efficiently without manual buffer tuning.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.