Oossa

llama.cpp release b11494 adds CUDA warp optimizations

The open‑source Llama.cpp library now uses four GDN state columns per CUDA warp, improving GPU performance.

NoteBy Published by Oossa: 1 min read

The ggml‑org team released llama.cpp version b11494. The update changes the CUDA backend to assign four GDN state columns per warp instead of two. This tweak is the default setting now. Binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds are available for download.

Why it matters

Developers using llama.cpp on NVIDIA GPUs can expect faster inference without changing their code.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.