A walkthrough of MBRL: Dyna, MCTS and the AlphaGo line
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
I've been reading around model-based RL for a few months and ended up writing a long walkthrough. Part 1 is up; part 2 (the optimal-control side) is still being written. What I'd most like feedback on is the organizing frame (big picture), also the delivery and diagram. Link: https://medium.com/@mr…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-08-22 19:11 · r/reinforcementlearning
A walkthrough of MBRL: Dyna, MCTS and the AlphaGo line