Oossa

llama.cpp b11510 adds CUDA support for larger tensors

The open‑source LLaMA inference library releases version b11510 with a new CUDA kernel that handles more than 65,535 rows, plus binaries for macOS, iOS and Linux.

NoteBy Published by Oossa: 1 min read

The ggml‑org team posted llama.cpp version b11510 on Oct 8 2026. It adds a looped PAD kernel for CUDA so the library can process tensors with over 65,000 rows or slices. The release also bundles pre‑built binaries for Apple Silicon, Intel macOS, iOS, and several Ubuntu builds (CPU, Vulkan and CUDA 12.8). The code and attestations are linked from the GitHub page.

Why it matters

Developers can now run larger language models on NVIDIA GPUs without hitting the previous row limit.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.