Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term
This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.
llmrun.dev estimates whether a given LLM can be run locally and how fast it will decode. The first part of that question is arithmetic. The second required several iterations, and Mixture-of-Experts models were the case where the initial formula was least accurate. This note describes the estimator…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-05 14:33 · DEV Community — Machine Learning
Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term