Oossa

llama.cpp adds new buffer allocation API in b11351 release

The Oct 2 2026 update introduces alloc_buffer_n to ggml’s buffer interface, improving multi‑device memory handling.

NoteOossaPublished by Oossa: 1 min read

The ggml‑org/llama.cpp project released version b11351 on 2026‑10‑02. The update adds an alloc_buffer_n method to the ggml_backend_buffer_type_i interface, with a public API called ggml_backend_buft_alloc_buffer_n. The default implementation now supports splitting buffers across multiple devices and better tensor allocation. Existing buffer types like CPU, Metal, OpenVINO and others get a NULL placeholder for the new method. Tests were added to cover the new functionality.

Why it matters

Developers can now manage memory more reliably on diverse hardware, reducing allocation errors in AI models.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.