Researchers introduced CureWM, a technique that adds simulated failure examples to existing robot world models. It builds alternative actions from successful demos, checks their outcomes in simulation or on a real arm, and fine‑tunes the model on these verified failures. On 484 held‑out LIBERO tasks, optimism dropped from 80% to 30‑43% across four fine‑tuned models. Physical robot tests showed false‑success scores fall from 90% to 33%.
Why it matters
Reduced false‑success predictions help robot planners avoid unsafe actions in real‑world deployments.