I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens
First of all, thank you for reading. I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons. Apex-2 - Architecture: Decoder-only MoE, every layer is MoE (no dense layers) - Size: 3.87B total parameters, 1.45B active per t…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-10-04 15:53 · r/huggingface
I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens - 2026-10-04 15:50 · r/LocalLLaMA
I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens