Qwen3.8-Flash-Next (104 GB MoE) on a Strix Halo + RTX 3090 Ti eGPU: 22 -> 84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time" title="Qwen3.8-Flash-Next (104 GB MoE) on a Strix Halo + RTX 3090 Ti eGPU: 22 -> 84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time" />
Read the full story at r/LocalLLM ↗