ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing
arXiv:2609.38541v1 Announce Type: new Abstract: Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic r…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CV
ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing