Oossa

llama.cpp adds faster CUDA top‑k selection

The open‑source Llama.cpp library now uses a radix‑select algorithm for CUDA top‑k, cutting runtime from 5.76 s to 0.94 s on a large test.

NoteBy Published by Oossa: 1 min read

The llama.cpp project released a new version (b11513) on Oct 8 2026. It replaces the previous per‑row CUDA top‑k routine with a grid‑over‑rows radix select. On the Qwen‑4‑exp model at 34,816 tokens, the change reduces top‑k calls from 1,671,253 launches (5,761.8 ms) to 2,329 launches (941.8 ms). The update also adds automatic selection of the best top‑k algorithm based on matrix shape and refines thresholds for bitonic and radix paths.

Why it matters

Developers running large language models on Nvidia GPUs can expect roughly six‑fold faster top‑k operations, speeding up inference and training.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.