The llama.cpp project published a new pre‑release called b11517 on 9 Oct 2026. The update adds a CUDA option that lets users pass the source‑1 precision to the MMQ helpers, a low‑level part of the GPU code. The change is signed by a NVIDIA contributor and applies to the CUDA 12 and 13 builds listed for Linux and Windows. The release also includes build targets for many other platforms, from macOS Apple Silicon to Android Snapdragon.
The update is available from the project's GitHub page and the llama.app website.
Why it matters
Developers can now control GPU precision in llama.cpp, which may improve performance or memory use for running LLaMA models.