Automated Reinforcement Learning should scare you
An LLM's training can be roughly divided into two stages: supervised learning (SL) and reinforcement learning (RL). In SL, you curate a dataset of text and train the LLM to predict the next token in that text from the tokens before it. The goal at this stage is to produce a model which is capable o…
Read the full story at r/artificial ↗
Timeline · 1 report
- 2026-09-21 18:52 · r/artificial
Automated Reinforcement Learning should scare you