# Mid‑training Qwen3 on raw web video boosts vision benchmarks

> Researchers found that feeding a 1.7 B‑parameter language model raw video clips improves its image and video task scores without hurting text ability.

Oossa · 2026-10-09 · https://oossa.com/en/mid-training-qwen3-on-raw-web-video-boosts-vision-benchmarks

Jaedong Hwang and colleagues mid‑trained the Qwen3‑1.7B model on uncaptioned clips from the YT‑Temporal‑1B video set. After the same instruction tuning as a baseline model, the video‑mid‑trained version scored 2.9 points higher on average across four video benchmarks and 5.1 points higher on ten image benchmarks. Text performance stayed essentially unchanged, with a 48.9 average versus 48.0 for the baseline.

## The facts

- Mid‑training used raw clips from YT‑Temporal‑1B
- Published on 2026‑10‑09

## Why it matters

The result shows that large language models can learn visual skills from raw video without any captions, opening a cheaper path to better multimodal AI.

## Sources & references

1. [Mid-Training Language Models on Raw Video](https://arxiv.org/abs/2610.11019) – arXiv cs.CV, 2026-10-09

Last updated: 2026-10-09
