# llama.cpp adds new buffer allocation API in b11351 release

> The Oct 2 2026 update introduces alloc_buffer_n to ggml’s buffer interface, improving multi‑device memory handling.

Oossa · 2026-10-02 · https://oossa.com/en/llama-cpp-adds-new-buffer-allocation-api-in-b11351-release

The ggml‑org/llama.cpp project released version b11351 on 2026‑10‑02. The update adds an alloc_buffer_n method to the ggml_backend_buffer_type_i interface, with a public API called ggml_backend_buft_alloc_buffer_n. The default implementation now supports splitting buffers across multiple devices and better tensor allocation. Existing buffer types like CPU, Metal, OpenVINO and others get a NULL placeholder for the new method. Tests were added to cover the new functionality.

## The facts

- Version b11351 released on 2026‑10‑02
- New alloc_buffer_n method added to ggml buffer interface

## Why it matters

Developers can now manage memory more reliably on diverse hardware, reducing allocation errors in AI models.

## Sources & references

1. [ggml-org/llama.cpp b11351](https://github.com/ggml-org/llama.cpp/releases/tag/b11351) – llama.cpp, 2026-10-02

Last updated: 2026-10-02
