The llama.cpp project released version b11509 on Oct 8, 2026. It fixes a CUDA guard that mis‑identified CCCL version 4.x as older, causing a silent performance drop in the argsort operation. The fix replaces the component‑wise check with a single packed integer comparison, restoring the fast strided‑iterator path. The change was tested on an RTX 4070 with CUDA 13.4 and passed all backend tests.
Why it matters
Developers using llama.cpp on modern CUDA setups will see faster sorting performance without changing their code.