Researchers introduced Robo‑COP, a method where a robot’s vision‑language‑action policy and its orchestrator are fine‑tuned together during real use. The system gathers its own mistake data, updates the policy, and only adopts the new version after it proves better on the tasks it was trained for. In ten simulated RoboLab tasks, success rose from 64.8% to 73.8%, and on three real‑world tasks it climbed from 38.3% to 50.0%.
Why it matters
It lets service robots get better on the job, cutting the need for frequent manual reprogramming.