I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
My goal was hands-on experience training a flow model from scratch, not just fine-tuning someone else's. So I built and trained one: a 210M-parameter diffusion transformer, 4.2M curated images at 256², rectified flow on the FLUX.2 VAE, flan-t5-base for text (128 tokens max). Only those two frozen p…
Read the full story at r/StableDiffusion ↗
Timeline · 2 reports
- 2026-09-11 13:00 · r/MachineLearning
Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P] - 2026-09-10 13:18 · r/StableDiffusion
I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days