# Google reports 2.4× faster sparse video attention on TPUs

> Google says a tile-aligned attention kernel cut latency on one TPU v6e chip. The reported speedup covers the attention kernel, not the full video-generation pipeline.

Oossa · 2026-09-30 · https://oossa.com/en/google-reports-2-4-faster-sparse-video-attention-on-tpus

Event date: 2026-09-30

Produced and translated with AI assistance. Check the original sources below.

Google published a case study on speeding up video-diffusion attention, a costly step in generating long, high-resolution clips. Its JAX and Pallas kernel uses sparse attention and aligns the attention mask with the tiles the TPU computes.

In tests on one TPU v6e chip, the final version took 32.76 milliseconds, versus 78.70 milliseconds for dense Splash attention—a reported 2.40× speedup. The measurements used synthetic inputs and exclude routing, token rearrangement and communication between devices, so they do not establish an end-to-end video-generation speedup.

## The facts

- The tests used 75,600 tokens, 10 heads and a head dimension of 128 on one TPU v6e chip.
- Google reported 32.76 ms for its final sparse kernel and 78.70 ms for dense Splash attention.

## Why it matters

For teams optimizing video generation on TPUs, the result shows that a sparse mask can be slower than dense attention unless the kernel is designed to skip and efficiently process the right tiles; the source does not show the impact on full video-generation time.

## Sources & references

1. [Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs](https://developers.googleblog.com/en/accelerating-spatio-temporal-attention-for-video-diffusion-on-tpus/) – Google for Developers

Last updated: 2026-09-30
