AMD Instella-MoE: How a Fully Open 16B MoE Model Proves You Don't Need NVIDIA to Train at Scale
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
AMD Instella-MoE: How a Fully Open 16B MoE Model Proves You Don't Need NVIDIA to Train at Scale When AMD released Instella-MoE-16B-A3B in September 2026, it did something that most large model releases don't: it published everything. Not just the weights, but the training code, data mixtures, confi…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-09 16:04 · DEV Community — Machine Learning
AMD Instella-MoE: How a Fully Open 16B MoE Model Proves You Don't Need NVIDIA to Train at Scale