"Scaling Post-Training Is All We Did": RL Environments Are Now a Data-Supply Problem
When Z.ai shipped GLM-5.3 this month, one line in the announcement got quoted everywhere: scaling post-training is all we did. No new base architecture, no bigger pretraining run. The gains came from reinforcement learning on a wider set of task environments. That sentence is a decent summary of Se…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-29 02:07 · DEV Community — Machine Learning
"Scaling Post-Training Is All We Did": RL Environments Are Now a Data-Supply Problem