Oossa

llama.cpp update speeds GLM prompts on Intel Arc GPU

The new SYCL changes boost token throughput up to 3× on an Intel Arc Pro B70, and reduce memory traffic for flash attention.

NoteBy Published by Oossa: 1 min read

The llama.cpp library added SYCL support that accelerates GLM‑4.7 flash attention. On an Intel Arc Pro B70 GPU the prompt‑prefill speed for a 8192‑token context rose from 432.80 to 1,292.29 tokens per second – almost three times faster. A second tweak stores intermediate scores in half‑precision (F16) instead of single‑precision, shaving a few percent off processing time while cutting memory use. Other workloads stayed the same.

Why it matters

Developers can run larger GLM models faster on affordable Intel Arc GPUs, making interactive AI apps more responsive.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.