Hanwen Li and colleagues released a paper on Oct 8 2026 describing VoS (Value of Steering). It learns from a large table of counterfactual runs to decide at which step an LLM‑based agent should be corrected. VoS can run offline or online and uses a harm‑budget trigger to avoid disturbing successful runs. Across 12 benchmark‑agent combinations, VoS raised scores by an average of 7.8 points and beat five existing uncertainty‑based triggers in 11 cases.
Why it matters
For developers of AI assistants, VoS offers a practical way to intervene only when it’s likely to help, improving reliability without excessive manual oversight.