AINewsnow

I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs β€” 500+ languages, Apache 2.0

I spent the last few months trying something that probably isn't the most practical way to build a translation model πŸ˜… I trained **Mythos 2.0**, a 2.04B-parameter Sparse Mixture-of-Experts translation model from scratch, using NVIDIA RTX 5090 GPUs. ### Key Highlights: * 🌍 **500+ languages** suppo…

Read the full story at r/machinelearningnews β†—

Timeline Β· 2 reports

  1. 2026-09-20 13:55 Β· r/deeplearning
    I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs β€” 500+ languages, Apache 2.0
  2. 2026-09-20 13:43 Β· r/machinelearningnews
    I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs β€” 500+ languages, Apache 2.0

More stories

  1. NVIDIA CEO Jensen Huang rejects β€˜AI will end the world’ claim, yet cautions β€˜we should go as fast as we can but...’ β€” Mint AI
  2. 5 Companies Using NVIDIA AI for Clean Energy β€” NVIDIA Blog
  3. On theCUBE Pod: Dreamforce unveils the agentic dream, and Nscale goes public β€” SiliconANGLE AI
  4. Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. β€” r/LocalLLaMA
  5. FREE AI TRAINING CREDIT β€” r/learnmachinelearning
  6. Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark β€” r/LocalLLaMA
  7. Nvidia chip-use emissions match that of a Russian state coal producer, report says β€” Euronews Next
  8. No one is surprised that Nvidia’s Jensen Huang thinks AI fears are overblown β€” The Verge AI

Get the daily brief of stories like this at 6:30 every morning β†’