Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation
arXiv:2609.25048v1 Announce Type: new Abstract: How many prompts does on-policy distillation (OPD) need, and how does the answer depend on the student policies that generate its training responses? We study these two controls jointly: prompt breadth and rollout refresh. A 3x3 mathematical-reasoning…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-23 04:00 · arXiv cs.CL
Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation