Training an Interactive World Model from Scratch (Actions -> Video Diffusion)
I've been wanting to understand interactive video generation / world models a bit better, so I decided to try training a small one from scratch. I generated ~30 hours of simple gameplay data and trained an action-conditioned diffusion transformer to predict the next frame given the previous frames…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-25 02:41 · r/learnmachinelearning
Training an Interactive World Model from Scratch (Actions -> Video Diffusion)