Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0)
Hey everyone! For the past few months, I've been training a 2.04B Sparse Mixture-of-Experts (SMoE) foundation translation model (Mythos2.0-2B) from scratch on 2x T4 GPUs. It covers 500+ languages—focusing on underserved African, Indigenous, and regional Asian languages that have zero commercial API…
Read the full story at r/huggingface ↗
Timeline · 3 reports
- 2026-09-20 13:55 · r/deeplearning
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0 - 2026-09-20 13:43 · r/machinelearningnews
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0 - 2026-09-18 15:46 · r/huggingface
Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0)