The ggml‑org/llama.cpp project released version b11351 on 2026‑10‑02. The update adds an alloc_buffer_n method to the ggml_backend_buffer_type_i interface, with a public API called ggml_backend_buft_alloc_buffer_n. The default implementation now supports splitting buffers across multiple devices and better tensor allocation. Existing buffer types like CPU, Metal, OpenVINO and others get a NULL placeholder for the new method. Tests were added to cover the new functionality.
Why it matters
Developers can now manage memory more reliably on diverse hardware, reducing allocation errors in AI models.