# Tencent releases Youtu-Parsing-Omni, a 5B omni‑modal parsing model

> The new model on Hugging Face can read documents, images, charts, audio and video, outputting a single JSON structure.

Oossa · 2026-10-09 · https://oossa.com/en/tencent-releases-youtu-parsing-omni-a-5b-omni-modal-parsing-model

Tencent has uploaded Youtu-Parsing-Omni to Hugging Face. It is a compact 5 billion‑parameter model that can take a single input—such as a page of text, a photograph, a chart, an audio clip or a video—and return a structured JSON envelope describing layout, text, tables, formulas, timestamps and captions. The model achieved the highest Overall score on the OmniDocBench benchmark among open models and ranked second only to Gemini‑3‑Pro on OmniParsingBench. A vLLM plugin and inference examples are included for easy serving.

## The facts

- 5 B parameters
- Released on Hugging Face in October 2026

## Why it matters

Developers can add multi‑type document understanding to apps without needing separate OCR, audio or video models.

## Sources & references

1. [New model on Hugging Face: tencent/Youtu-Parsing-Omni (image-text-to-text)](https://huggingface.co/tencent/Youtu-Parsing-Omni) – Tencent Hunyuan, 2026-10-09

Last updated: 2026-10-09
