Hot Expert Reload on GPU is what this community needs
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on MOE models with moderate number of active parameters, like Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, GLM 5.3 Flash. With 2x 3090 speeds will be qui…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-11 16:25 · r/LocalLLaMA
Hot Expert Reload on GPU is what this community needs