The ggml‑org/llama.cpp project released version b11372 on Oct 3 2026. The update rewrites the Qwen4exp indexer so each head computes its score separately and writes results in place. This halves the memory needed for the indexer score buffers, which were the biggest memory users in long‑context graphs. The change does not affect compute speed and also adds support for the new lightning indexer on CUDA, Metal and Vulkan back‑ends.
Why it matters
Developers running large language models with long context will be able to fit bigger prompts on the same hardware.