Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
arXiv:2609.28851v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the re…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CV
Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning