Tencent has uploaded Youtu-Parsing-Omni to Hugging Face. It is a compact 5 billion‑parameter model that can take a single input—such as a page of text, a photograph, a chart, an audio clip or a video—and return a structured JSON envelope describing layout, text, tables, formulas, timestamps and captions. The model achieved the highest Overall score on the OmniDocBench benchmark among open models and ranked second only to Gemini‑3‑Pro on OmniParsingBench. A vLLM plugin and inference examples are included for easy serving.
Why it matters
Developers can add multi‑type document understanding to apps without needing separate OCR, audio or video models.