Oossa

llama.cpp pre‑release b11517 adds CUDA precision option

The open‑source LLaMA inference library released version b11517 on Oct 9, 2026, adding a CUDA flag to set source‑1 precision.

NoteBy Published by Oossa: 1 min read

The llama.cpp project published a new pre‑release called b11517 on 9 Oct 2026. The update adds a CUDA option that lets users pass the source‑1 precision to the MMQ helpers, a low‑level part of the GPU code. The change is signed by a NVIDIA contributor and applies to the CUDA 12 and 13 builds listed for Linux and Windows. The release also includes build targets for many other platforms, from macOS Apple Silicon to Android Snapdragon.

The update is available from the project's GitHub page and the llama.app website.

Why it matters

Developers can now control GPU precision in llama.cpp, which may improve performance or memory use for running LLaMA models.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.