The ggml‑org team published llama.cpp release b11531. It refactors the chat API and bundles pre‑built binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu variants including CPU, Vulkan and CUDA 12.8. The download links are on the GitHub release page. Users can now run LLMs locally on more platforms without compiling themselves.
Why it matters
Developers can run Llama models locally on more devices today, avoiding build steps and supporting GPU acceleration on Ubuntu.