Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models
arXiv:2609.22293v1 Announce Type: new Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.CV
Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models