# WAM-Cache cuts robot vision compute by up to 42%

> The new training‑free cache reduces video model cost while keeping manipulation accuracy within a few points.

Oossa · 2026-10-09 · https://oossa.com/en/wam-cache-cuts-robot-vision-compute-by-up-to-42

Researchers led by Kai Ding released WAM-Cache, a framework that reuses key‑value representations from a video diffusion transformer across control chunks. Instead of recomputing the whole visual encoder each step, the system refreshes only a sparse set of tokens chosen by the action expert’s attention and a visual surprise measure, with a strict age limit. On the Fast‑WAM benchmark it lowers video encoder FLOPs by 32‑42% on RoboTwin 2.0, LIBERO and real‑world tests, while staying within 0.7‑1.8 percentage points of the dense baseline in simulation and 2.5 points on a real robot.

## The facts

- WAM-Cache reduces video DiT prefill FLOPs by 32‑42% (Fast‑WAM benchmark).
- Accuracy loss is at most 1.8 % in simulation and 2.5 % on a real robot.

## Why it matters

Robotics developers can run high‑performance manipulation policies on cheaper hardware without a large drop in success rates.

## Sources & references

1. [WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models](https://arxiv.org/abs/2610.11401) – arXiv cs.RO, 2026-10-09

Last updated: 2026-10-09
