# llama.cpp release b11494 adds CUDA warp optimizations

> The open‑source Llama.cpp library now uses four GDN state columns per CUDA warp, improving GPU performance.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-release-b11494-adds-cuda-warp-optimizations

The ggml‑org team released llama.cpp version b11494. The update changes the CUDA backend to assign four GDN state columns per warp instead of two. This tweak is the default setting now. Binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds are available for download.

## The facts

- Release tag b11494 published on 2026-10-08
- CUDA default cols_per_warp set to 4

## Why it matters

Developers using llama.cpp on NVIDIA GPUs can expect faster inference without changing their code.

## Sources & references

1. [ggml-org/llama.cpp b11494](https://github.com/ggml-org/llama.cpp/releases/tag/b11494) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
