# llama.cpp b11494 optimiert CUDA-Warps

> Die Open-Source-Bibliothek llama.cpp nutzt jetzt vier GDN-State-Spalten pro CUDA-Warp und verbessert damit die GPU-Leistung.

Oossa · 2026-10-08 · https://oossa.com/de/llama-cpp-release-b11494-adds-cuda-warp-optimizations

Das ggml-org-Team hat llama.cpp in Version b11494 veröffentlicht. Das Update ändert das CUDA-Backend: Es weist jedem Warp nun vier statt zwei GDN-State-Spalten zu. Diese Einstellung ist jetzt der Standard. Binaries für macOS (Apple Silicon und Intel), iOS und verschiedene Ubuntu-Versionen stehen zum Download bereit.

## Die Fakten

- Release-Tag b11494 am 2026-10-08 veröffentlicht
- CUDA-Standardwert für cols_per_warp auf 4 gesetzt

## Warum es wichtig ist

Entwickler, die llama.cpp auf NVIDIA-GPUs nutzen, können mit schnellerer Inferenz rechnen, ohne ihren Code ändern zu müssen.

## Quellen und Referenzen

1. [ggml-org/llama.cpp b11494](https://github.com/ggml-org/llama.cpp/releases/tag/b11494) – llama.cpp, 2026-10-08

Zuletzt aktualisiert: 2026-10-08
