From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
arXiv:2609.30391v1 Announce Type: new Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: w…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv cs.LG
From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning