Less from More: Reinforcing Sparse Video Reasoning from Dense References
arXiv:2610.10893v1 Announce Type: new Abstract: Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CV
Less from More: Reinforcing Sparse Video Reasoning from Dense References