# llama.cpp pre‑release b11517 adds CUDA precision option

> The open‑source LLaMA inference library released version b11517 on Oct 9, 2026, adding a CUDA flag to set source‑1 precision.

Oossa · 2026-10-09 · https://oossa.com/en/llama-cpp-pre-release-b11517-adds-cuda-precision-option

The llama.cpp project published a new pre‑release called b11517 on 9 Oct 2026. The update adds a CUDA option that lets users pass the source‑1 precision to the MMQ helpers, a low‑level part of the GPU code. The change is signed by a NVIDIA contributor and applies to the CUDA 12 and 13 builds listed for Linux and Windows. The release also includes build targets for many other platforms, from macOS Apple Silicon to Android Snapdragon.

The update is available from the project's GitHub page and the llama.app website.

## The facts

- Release tag b11517 published on 09 Oct 2026
- Adds CUDA flag for source‑1 precision (affects CUDA 12 and 13 builds)

## Why it matters

Developers can now control GPU precision in llama.cpp, which may improve performance or memory use for running LLaMA models.

## Sources & references

1. [ggml-org/llama.cpp b11517](https://github.com/ggml-org/llama.cpp/releases/tag/b11517) – llama.cpp, 2026-10-09

Last updated: 2026-10-09
