Mid-Training Language Models on Raw Video
arXiv:2610.11019v1 Announce Type: new Abstract: Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CV
Mid-Training Language Models on Raw Video