# llama.cpp v b11499 adds CUDA roll improvement and new binaries

> The open‑source LLM runtime llama.cpp released version b11499 on Oct 8 2026, adding a CUDA fix for non‑contiguous roll ops and new pre‑built binaries for macOS, iOS and Linux.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-v-b11499-adds-cuda-roll-improvement-and-new-binaries

The ggml‑org team posted llama.cpp release b11499 on Oct 8 2026. The update changes the CUDA backend to use byte strides for the roll operation, which lets the GPU handle non‑contiguous data correctly. The release also bundles fresh binaries for Apple Silicon, Intel macOS, iOS, and several Ubuntu builds (CPU, Vulkan and CUDA 12). All files are available for download from the GitHub release page.

## The facts

- Release version b11499 published on 2026‑10‑08
- CUDA backend now supports byte‑stride roll operations

## Why it matters

Developers can run llama.cpp on more GPU setups without errors, expanding local LLM use on desktops and servers.

## Sources & references

1. [ggml-org/llama.cpp b11499](https://github.com/ggml-org/llama.cpp/releases/tag/b11499) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
