The ggml‑org team published llama.cpp b11496 on Oct 8 2026. The update adds a SYCL stage that moves bulk model uploads through a pinned ring buffer, speeding up loading on compatible GPUs. Binaries for macOS (Apple Silicon and Intel), iOS, and various Ubuntu builds (CPU, Vulkan, CUDA 12) are available for download.
Why it matters
Developers can load large language models faster on SYCL‑compatible hardware, reducing startup time for apps that use llama.cpp.