Oossa

Mid‑training Qwen3 on raw web video boosts vision benchmarks

Researchers found that feeding a 1.7 B‑parameter language model raw video clips improves its image and video task scores without hurting text ability.

NoteBy Published by Oossa: 1 min read

Jaedong Hwang and colleagues mid‑trained the Qwen3‑1.7B model on uncaptioned clips from the YT‑Temporal‑1B video set. After the same instruction tuning as a baseline model, the video‑mid‑trained version scored 2.9 points higher on average across four video benchmarks and 5.1 points higher on ten image benchmarks. Text performance stayed essentially unchanged, with a 48.9 average versus 48.0 for the baseline.

Why it matters

The result shows that large language models can learn visual skills from raw video without any captions, opening a cheaper path to better multimodal AI.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.