The llama.cpp library added SYCL support that accelerates GLM‑4.7 flash attention. On an Intel Arc Pro B70 GPU the prompt‑prefill speed for a 8192‑token context rose from 432.80 to 1,292.29 tokens per second – almost three times faster. A second tweak stores intermediate scores in half‑precision (F16) instead of single‑precision, shaving a few percent off processing time while cutting memory use. Other workloads stayed the same.
Why it matters
Developers can run larger GLM models faster on affordable Intel Arc GPUs, making interactive AI apps more responsive.