tencent/Youtu-Parsing-Omni · Hugging Face
Youtu-Parsing-Omni is a compact (5B) omni-modal parsing model. Given a single input — a document page, a natural image, a chart / flowchart, a geometry figure, an audio clip or an audio-visual video — it produces one structured JSON envelope that covers both perception (layout elements, text, table…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-09 12:03 · r/LocalLLaMA
tencent/Youtu-Parsing-Omni · Hugging Face