# llama.cpp adds 2‑aligned F32 loads for Intel Vulkan builds

> The update improves matrix multiplication speed on Intel GPUs by loading two floats at a time.

Oossa · 2026-09-30 · https://oossa.com/en/llama-cpp-adds-2-aligned-f32-loads-for-intel-vulkan-builds

The llama.cpp project released a patch that changes how Vulkan code loads 32‑bit floats (F32). On Intel GPUs, loading two floats together is faster than one at a time, so the new logic uses a 2‑aligned load in the mul_mat_vec routine. The change also adds a missing alignment check. Benchmarks on a B60 test board show up to 1.63 TFLOPS for certain matrix sizes, a noticeable gain over the previous version.

## The facts

- Patch released Sep 30, 2026 (b11266)
- Speedup to 1.63 TFLOPS on a 4096×8×14336 matrix

## Why it matters

Faster GPU math means quicker LLM inference for users running llama.cpp on Intel hardware.

## Sources & references

1. [ggml-org/llama.cpp b11266](https://github.com/ggml-org/llama.cpp/releases/tag/b11266) – llama.cpp, 2026-09-29
2. [ggml-org/llama.cpp b11267](https://github.com/ggml-org/llama.cpp/releases/tag/b11267) – llama.cpp, 2026-09-30

Last updated: 2026-09-30
