Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
Coverage of "Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!" from 2 sources, with a live timeline of who reported what and when.
Read the full story at r/learnmachinelearning ↗
Timeline · 2 reports
- 2026-09-06 18:41 · r/MachineLearning
Proposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p] - 2026-09-05 07:16 · r/learnmachinelearning
Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!