AINewsnow

I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

My goal was hands-on experience training a flow model from scratch, not just fine-tuning someone else's. So I built and trained one: a 210M-parameter diffusion transformer, 4.2M curated images at 256², rectified flow on the FLUX.2 VAE, flan-t5-base for text (128 tokens max). Only those two frozen p…

Read the full story at r/StableDiffusion ↗

Timeline · 2 reports

  1. 2026-09-11 13:00 · r/MachineLearning
    Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P]
  2. 2026-09-10 13:18 · r/StableDiffusion
    I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days

More stories

  1. The Local Dream 3 is coming with some huge new features! — r/StableDiffusion
  2. Need a bit of help with flux 2 klein image generation to follow style+palete from reference. — r/comfyui
  3. Is there a way to reverse image search with models to see what it recognizes? — r/StableDiffusion
  4. Please teach me how to image to image — r/comfyui
  5. Need help with LTX 2.3 — r/StableDiffusion
  6. Flux klein 4b lora training — r/comfyui
  7. Minimax H3+ Flux 2 pro japan bicycle street video workflow — r/StableDiffusion
  8. Best model/workflow for face and body consistency? — r/comfyui

Get the daily brief of stories like this at 6:30 every morning →