CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these practical constraints, we study whether text…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-17 04:00 · arXiv cs.AI
CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video