AINewsnow

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

First of all, thank you for reading. I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons. Apex-2 - Architecture: Decoder-only MoE, every layer is MoE (no dense layers) - Size: 3.87B total parameters, 1.45B active per t…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-04 15:53 · r/huggingface
    I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens
  2. 2026-10-04 15:50 · r/LocalLLaMA
    I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. An OpenAI safety employee has quit and is sounding the alarm — The Verge AI
  3. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  4. A model guide for the GPT-6 family — OpenAI News
  5. OpenAI fires 3 AI safety researchers for allegedly sharing confidential company information — Mint AI
  6. Introducing Oscilloscope Diffusion — r/comfyui
  7. Apple says it's tightening macOS Full Disk Access' controls due to new risks from AI agents — TechCrunch AI
  8. Google launches satellite to test feasibility of building data centers in space — NPR Technology

Get the daily brief of stories like this at 6:30 every morning →