Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges
arXiv:2609.31665v1 Announce Type: new Abstract: Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different st…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.CV
Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges