"The Talker Does Not Control The Doer (in Current AIs)", Eliezer Yudkowsky (on 'chunky posttraining' and RL misgeneralization/reward-hacking in frontier LLMs like Astra/Mythos)
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
Coverage of ""The Talker Does Not Control The Doer (in Current AIs)", Eliezer Yudkowsky (on 'chunky posttraining' and RL misgeneralization/reward-hacking in frontier LLMs like Astra/Mythos)" from 1 source, with a live timeline of who reported what and when.
Read the full story at r/reinforcementlearning ↗