The llama.cpp project released version b11490 on 2026‑10‑08. It adds support for the Hexagon "alloc_buffer_n" API, allowing large tensors to be split across multiple buffers. The default dynamic buffer size is increased from 128 MB to 512 MB to avoid performance drops with big models. A new "--no-embd-offload" flag simplifies command lines for devices that cannot offload embeddings. The changes were co‑authored by Jhen‑Jie Hong.
Why it matters
Developers using Hexagon‑based hardware can run larger language models more efficiently without manual buffer tuning.