# llama.cpp adds faster CUDA top‑k selection

> The open‑source Llama.cpp library now uses a radix‑select algorithm for CUDA top‑k, cutting runtime from 5.76 s to 0.94 s on a large test.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-adds-faster-cuda-top-k-selection

The llama.cpp project released a new version (b11513) on Oct 8 2026. It replaces the previous per‑row CUDA top‑k routine with a grid‑over‑rows radix select. On the Qwen‑4‑exp model at 34,816 tokens, the change reduces top‑k calls from 1,671,253 launches (5,761.8 ms) to 2,329 launches (941.8 ms). The update also adds automatic selection of the best top‑k algorithm based on matrix shape and refines thresholds for bitonic and radix paths.

## The facts

- Release date: Oct 8 2026 – "PUBLISHED: Thu Oct 08 2026"
- Runtime drop on Qwen‑4‑exp: from 5,761.8 ms to 941.8 ms

## Why it matters

Developers running large language models on Nvidia GPUs can expect roughly six‑fold faster top‑k operations, speeding up inference and training.

## Sources & references

1. [ggml-org/llama.cpp b11513](https://github.com/ggml-org/llama.cpp/releases/tag/b11513) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
