Oossa

llama.cpp v0.2.4 adds SYCL Q5_K layout and GLU fusion

The open‑source llama.cpp library released version b11515 on Oct 9 2026 with new SYCL support and pre‑built binaries for macOS, iOS, Linux and CUDA.

NoteBy Published by Oossa: 1 min read

The ggml‑org team posted llama.cpp release b11515 on 2026‑10‑09. The update adds a SYCL‑based Q5_K reorder‑layout matrix‑multiply‑vector‑quantization (MMVQ) and a fused GLU operator, improving performance on compatible hardware. The release also bundles ready‑to‑run binaries for Apple Silicon, Intel Macs, iOS, Ubuntu (CPU, Vulkan, CUDA 12.8) and IBM s390x. Users can download the appropriate archive from the GitHub page.

Why it matters

Developers can now run llama.cpp faster on SYCL‑compatible GPUs without building from source.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.