When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
arXiv:2610.00202v1 Announce Type: new Abstract: Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correc…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.CL
When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora