Oossa

Google reports 2.4× faster sparse video attention on TPUs

Google says a tile-aligned attention kernel cut latency on one TPU v6e chip. The reported speedup covers the attention kernel, not the full video-generation pipeline.

NoteOossaPublished by Oossa: 1 min read

Google published a case study on speeding up video-diffusion attention, a costly step in generating long, high-resolution clips. Its JAX and Pallas kernel uses sparse attention and aligns the attention mask with the tiles the TPU computes.

In tests on one TPU v6e chip, the final version took 32.76 milliseconds, versus 78.70 milliseconds for dense Splash attention—a reported 2.40× speedup. The measurements used synthetic inputs and exclude routing, token rearrangement and communication between devices, so they do not establish an end-to-end video-generation speedup.

Why it matters

For teams optimizing video generation on TPUs, the result shows that a sparse mask can be slower than dense attention unless the kernel is designed to skip and efficiently process the right tiles; the source does not show the impact on full video-generation time.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.