RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explic…
Read the full story at Apple Machine Learning Research ↗
Timeline · 1 report
- 2026-10-06 00:00 · Apple Machine Learning Research
RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation