AINewsnow

Fully asynchronous multi-turn RL while keeping the training side in Transformers

i work on Halo at White Circle. one problem we wanted to solve was running environment rollouts without turning the training stack into a second serving framework. in Halo, multi-turn rollouts run as Ray actors against vLLM or SGLang while optimization stays in a normal Transformers/TRL trainer. ge…

Read the full story at r/reinforcementlearning ↗

Timeline · 1 report

  1. 2026-09-21 17:43 · r/reinforcementlearning
    Fully asynchronous multi-turn RL while keeping the training side in Transformers

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Ahead of Sam Altman's UN address, OpenAI proposes new ways to track AI misalignment risks — Axios AI+
  4. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  5. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  6. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  7. Mathematicians Hate AI. They Can’t Quit It — Wired AI
  8. Python Workers are now generally available — Cloudflare Blog — AI

Get the daily brief of stories like this at 6:30 every morning →