# llama.cpp b11510 adds CUDA support for larger tensors

> The open‑source LLaMA inference library releases version b11510 with a new CUDA kernel that handles more than 65,535 rows, plus binaries for macOS, iOS and Linux.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-b11510-adds-cuda-support-for-larger-tensors

The ggml‑org team posted llama.cpp version b11510 on Oct 8 2026. It adds a looped PAD kernel for CUDA so the library can process tensors with over 65,000 rows or slices. The release also bundles pre‑built binaries for Apple Silicon, Intel macOS, iOS, and several Ubuntu builds (CPU, Vulkan and CUDA 12.8). The code and attestations are linked from the GitHub page.

## The facts

- Version b11510 released on 2026‑10‑08
- New CUDA PAD kernel supports >65535 rows

## Why it matters

Developers can now run larger language models on NVIDIA GPUs without hitting the previous row limit.

## Sources & references

1. [ggml-org/llama.cpp b11510](https://github.com/ggml-org/llama.cpp/releases/tag/b11510) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
