SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
arXiv:2610.10889v1 Announce Type: new Abstract: Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CV
SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages