The ggml‑org team posted llama.cpp version b11510 on Oct 8 2026. It adds a looped PAD kernel for CUDA so the library can process tensors with over 65,000 rows or slices. The release also bundles pre‑built binaries for Apple Silicon, Intel macOS, iOS, and several Ubuntu builds (CPU, Vulkan and CUDA 12.8). The code and attestations are linked from the GitHub page.
Why it matters
Developers can now run larger language models on NVIDIA GPUs without hitting the previous row limit.