Reasoning Instructions Can Break Answer Decoding in Vision--Language Models
arXiv:2609.29278v1 Announce Type: new Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B d…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CL
Reasoning Instructions Can Break Answer Decoding in Vision--Language Models