I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs β 500+ languages, Apache 2.0
I spent the last few months trying something that probably isn't the most practical way to build a translation model π I trained **Mythos 2.0**, a 2.04B-parameter Sparse Mixture-of-Experts translation model from scratch, using NVIDIA RTX 5090 GPUs. ### Key Highlights: * π **500+ languages** suppoβ¦
Read the full story at r/machinelearningnews β
Timeline Β· 2 reports
- 2026-09-20 13:55 Β· r/deeplearning
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs β 500+ languages, Apache 2.0 - 2026-09-20 13:43 Β· r/machinelearningnews
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs β 500+ languages, Apache 2.0