VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding
arXiv:2609.38413v1 Announce Type: new Abstract: Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped ar…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CV
VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding