The ggml‑org/llama.cpp project released version b11516 on 2026‑10‑09. The update fixes a bug in the Hexagon NPU path where the DMA ring could overflow, causing stale data and NaN/inf results. The code now drops the oldest descriptor when the ring is full and flushes the queue later. Test cases with large patch‑embed operations now pass on Hexagon hardware.
Why it matters
Developers using llama.cpp on Qualcomm Hexagon devices will get correct model outputs instead of erroneous NaNs.