AINewsnow

Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0)

Hey everyone! For the past few months, I've been training a 2.04B Sparse Mixture-of-Experts (SMoE) foundation translation model (Mythos2.0-2B) from scratch on 2x T4 GPUs. It covers 500+ languages—focusing on underserved African, Indigenous, and regional Asian languages that have zero commercial API…

Read the full story at r/huggingface ↗

Timeline · 3 reports

  1. 2026-09-20 13:55 · r/deeplearning
    I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0
  2. 2026-09-20 13:43 · r/machinelearningnews
    I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0
  3. 2026-09-18 15:46 · r/huggingface
    Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0)

More stories

  1. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  2. what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn? — r/LocalLLaMA
  3. Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters — FP4 performance of the Ascend 960PR doubles expectations — Tom's Hardware
  4. Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. — r/LocalLLaMA
  5. Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. — r/LocalLLaMA
  6. FREE AI TRAINING CREDIT — r/learnmachinelearning
  7. No one is surprised that Nvidia's Jensen Huang thinks AI fears are overblown. — The Verge AI
  8. Dario Says AI Should Slow Down. Jensen Wants to Go Full Steam Ahead. — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →