MoVISA: Multi-Token Reasoning for Video Object Segmentation
arXiv:2609.28956v1 Announce Type: new Abstract: Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we ob…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CV
MoVISA: Multi-Token Reasoning for Video Object Segmentation