Researchers led by Kai Ding released WAM-Cache, a framework that reuses key‑value representations from a video diffusion transformer across control chunks. Instead of recomputing the whole visual encoder each step, the system refreshes only a sparse set of tokens chosen by the action expert’s attention and a visual surprise measure, with a strict age limit. On the Fast‑WAM benchmark it lowers video encoder FLOPs by 32‑42% on RoboTwin 2.0, LIBERO and real‑world tests, while staying within 0.7‑1.8 percentage points of the dense baseline in simulation and 2.5 points on a real robot.
Why it matters
Robotics developers can run high‑performance manipulation policies on cheaper hardware without a large drop in success rates.