The llama.cpp project released a patch that changes how Vulkan code loads 32‑bit floats (F32). On Intel GPUs, loading two floats together is faster than one at a time, so the new logic uses a 2‑aligned load in the mul_mat_vec routine. The change also adds a missing alignment check. Benchmarks on a B60 test board show up to 1.63 TFLOPS for certain matrix sizes, a noticeable gain over the previous version.
Why it matters
Faster GPU math means quicker LLM inference for users running llama.cpp on Intel hardware.