What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study
arXiv:2610.03792v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language mode…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.CV
What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study