# llama.cpp 更新让 Intel Arc GPU 上的 GLM 提示词处理提速近 3 倍

> 新的 SYCL 改动让 Intel Arc Pro B70 的 token 处理速度最高提升 3 倍，并减少了 Flash Attention 的内存流量。

Oossa · 2026-10-07 · https://oossa.com/zh/llama-cpp-update-speeds-glm-prompts-on-intel-arc-gpu

llama.cpp 库新增了 SYCL 支持，可加速 GLM‑4.7 的 Flash Attention。在 Intel Arc Pro B70 GPU 上，处理 8192-token 上下文时的提示词预填充速度从每秒 432.80 个 token 提升至 1,292.29 个，接近原来的 3 倍。另一项改动将中间分数从单精度改为半精度（F16）存储，在减少内存占用的同时，让处理时间缩短了几个百分点。其他工作负载表现不变。

## 事实

- Intel Arc Pro B70 上，pp8192 的 token 处理速度从 432.80 tok/s 提升至 1,292.29 tok/s（2.99 倍）
- 以 F16 存储 Flash Attention 分数后，pp8192 的速度提升 4.6%（从 1,583.9 提升至 1,657.1 tok/s）

## 为什么重要

开发者可以在价格亲民的 Intel Arc GPU 上更快运行更大的 GLM 模型，让交互式 AI 应用响应更迅速。

## 来源与参考

1. [ggml-org/llama.cpp b11463](https://github.com/ggml-org/llama.cpp/releases/tag/b11463) – llama.cpp, 2026-10-07

最后更新: 2026-10-07
