# llama.cpp adds Hexagon buffer support and larger default allocation

> The open‑source llama.cpp library now lets Hexagon devices split large tensors and raises the default dynamic buffer to 512 MB.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-adds-hexagon-buffer-support-and-larger-default-allocation

The llama.cpp project released version b11490 on 2026‑10‑08. It adds support for the Hexagon "alloc_buffer_n" API, allowing large tensors to be split across multiple buffers. The default dynamic buffer size is increased from 128 MB to 512 MB to avoid performance drops with big models. A new "--no-embd-offload" flag simplifies command lines for devices that cannot offload embeddings. The changes were co‑authored by Jhen‑Jie Hong.

## The facts

- Release date: 2026‑10‑08 ("Thu Oct 08 2026 06:58:59 GMT+0200")
- Default dynamic buffer raised to 512 MB

## Why it matters

Developers using Hexagon‑based hardware can run larger language models more efficiently without manual buffer tuning.

## Sources & references

1. [ggml-org/llama.cpp b11490](https://github.com/ggml-org/llama.cpp/releases/tag/b11490) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
