# llama.cpp update speeds GLM prompts on Intel Arc GPU

> The new SYCL changes boost token throughput up to 3× on an Intel Arc Pro B70, and reduce memory traffic for flash attention.

Oossa · 2026-10-07 · https://oossa.com/en/llama-cpp-update-speeds-glm-prompts-on-intel-arc-gpu

The llama.cpp library added SYCL support that accelerates GLM‑4.7 flash attention. On an Intel Arc Pro B70 GPU the prompt‑prefill speed for a 8192‑token context rose from 432.80 to 1,292.29 tokens per second – almost three times faster. A second tweak stores intermediate scores in half‑precision (F16) instead of single‑precision, shaving a few percent off processing time while cutting memory use. Other workloads stayed the same.

## The facts

- pp8192 token rate improved from 432.80 to 1,292.29 tok/s (2.99×) on Intel Arc Pro B70
- Storing flash‑attention scores in F16 gave a 4.6% boost for pp8192 (1,583.9 → 1,657.1 tok/s)

## Why it matters

Developers can run larger GLM models faster on affordable Intel Arc GPUs, making interactive AI apps more responsive.

## Sources & references

1. [ggml-org/llama.cpp b11463](https://github.com/ggml-org/llama.cpp/releases/tag/b11463) – llama.cpp, 2026-10-07

Last updated: 2026-10-07
