Oossa

Tencent releases Youtu-Parsing-Omni, a 5B omni‑modal parsing model

The new model on Hugging Face can read documents, images, charts, audio and video, outputting a single JSON structure.

NoteBy Published by Oossa: 1 min read

Tencent has uploaded Youtu-Parsing-Omni to Hugging Face. It is a compact 5 billion‑parameter model that can take a single input—such as a page of text, a photograph, a chart, an audio clip or a video—and return a structured JSON envelope describing layout, text, tables, formulas, timestamps and captions. The model achieved the highest Overall score on the OmniDocBench benchmark among open models and ranked second only to Gemini‑3‑Pro on OmniParsingBench. A vLLM plugin and inference examples are included for easy serving.

Why it matters

Developers can add multi‑type document understanding to apps without needing separate OCR, audio or video models.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.