Oossa

WAM-Cache cuts robot vision compute by up to 42%

The new training‑free cache reduces video model cost while keeping manipulation accuracy within a few points.

NoteBy Published by Oossa: 1 min read

Researchers led by Kai Ding released WAM-Cache, a framework that reuses key‑value representations from a video diffusion transformer across control chunks. Instead of recomputing the whole visual encoder each step, the system refreshes only a sparse set of tokens chosen by the action expert’s attention and a visual surprise measure, with a strict age limit. On the Fast‑WAM benchmark it lowers video encoder FLOPs by 32‑42% on RoboTwin 2.0, LIBERO and real‑world tests, while staying within 0.7‑1.8 percentage points of the dense baseline in simulation and 2.5 points on a real robot.

Why it matters

Robotics developers can run high‑performance manipulation policies on cheaper hardware without a large drop in success rates.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.