The ggml‑org team released llama.cpp version b11494. The update changes the CUDA backend to assign four GDN state columns per warp instead of two. This tweak is the default setting now. Binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds are available for download.
Why it matters
Developers using llama.cpp on NVIDIA GPUs can expect faster inference without changing their code.