Qwen3.8-2.4T: The 2.4 Trillion Parameter Open-Weight MoE AI Model — Day 27/30
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
TL;DR — Qwen3.8-2.4T-A95B is a 2.4 trillion parameter mixture-of-experts model with 95B active parameters per token and a 1,048,576-token context window. Probes showed it nailing a coding task and a math word problem cleanly, but returning nothing at all on a simple JSON extraction after burning 70…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-17 14:16 · DEV Community — AI
Qwen3.8-2.4T: The 2.4 Trillion Parameter Open-Weight MoE AI Model — Day 27/30