OossaAI is evolving fast. We explain it simply.

llama.cpp adds 2‑aligned F32 loads for Intel Vulkan builds

The update improves matrix multiplication speed on Intel GPUs by loading two floats at a time.

NoteOossa1 min read

The llama.cpp project released a patch that changes how Vulkan code loads 32‑bit floats (F32). On Intel GPUs, loading two floats together is faster than one at a time, so the new logic uses a 2‑aligned load in the mul_mat_vec routine. The change also adds a missing alignment check. Benchmarks on a B60 test board show up to 1.63 TFLOPS for certain matrix sizes, a noticeable gain over the previous version.

Why it matters

Faster GPU math means quicker LLM inference for users running llama.cpp on Intel hardware.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Sources & references

#SourceOutletDateKey takeaway
1ggml-org/llama.cpp b11266 ↗llama.cppSep 29, 2026<details open> vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254) It turns out Intel doesn't particularly like loading F32s
2ggml-org/llama.cpp b11267 ↗llama.cppSep 30, 2026<details open> vulkan: Tune GDN kernel, fix Intel performance (#29476) * vulkan: tune GDN shader * tune for Intel </details> **Website:** -

2 sources

Last updated: ·Markdown·llms.txt