GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
arXiv:2609.38285v1 Announce Type: new Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies a…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CV
GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions