Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
arXiv:2610.00024v1 Announce Type: new Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.CV
Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null