The llama.cpp project released a new version (b11513) on Oct 8 2026. It replaces the previous per‑row CUDA top‑k routine with a grid‑over‑rows radix select. On the Qwen‑4‑exp model at 34,816 tokens, the change reduces top‑k calls from 1,671,253 launches (5,761.8 ms) to 2,329 launches (941.8 ms). The update also adds automatic selection of the best top‑k algorithm based on matrix shape and refines thresholds for bitonic and radix paths.
Why it matters
Developers running large language models on Nvidia GPUs can expect roughly six‑fold faster top‑k operations, speeding up inference and training.