Jaedong Hwang and colleagues mid‑trained the Qwen3‑1.7B model on uncaptioned clips from the YT‑Temporal‑1B video set. After the same instruction tuning as a baseline model, the video‑mid‑trained version scored 2.9 points higher on average across four video benchmarks and 5.1 points higher on ten image benchmarks. Text performance stayed essentially unchanged, with a 48.9 average versus 48.0 for the baseline.
Why it matters
The result shows that large language models can learn visual skills from raw video without any captions, opening a cheaper path to better multimodal AI.