I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0
This story is from 2026-09-20. It is preserved in the archive; the latest stories are on the live feed.
Coverage of "I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0" from 1 source, with a live timeline of who reported what and when.
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-20 13:55 · r/deeplearning
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0