What are the biggest open problems in long-video understanding right now?
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
I've been reading recent work on long-video understanding, particularly STORM: Token-Efficient Long Video Understanding for Multimodal LLMs. I'm trying to identify a research direction rather than just build another Video-LLM. I'm particularly interested in temporal modeling, video representations,…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-09 16:31 · r/deeplearning
What are the biggest open problems in long-video understanding right now?