Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P]
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I trained a 210M-parameter text-to-image diffusion transformer from scratch (3.5 days, one RTX PRO 6000, 4.2M images at 256²) mainly to understand the recipe end to end. Three measurements came out of it that I have not seen stated plainly elsewhere, so I'm posting those rather than the samples. 1.…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-11 13:00 · r/MachineLearning
Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P]