I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0
Coverage of "I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0" from 2 sources, with a live timeline of who reported what and when.
Read the full story at r/deeplearning ↗
Timeline · 2 reports
- 2026-09-26 02:44 · r/learnmachinelearning
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0 - 2026-09-26 02:40 · r/deeplearning
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0