Recipe & Patches: MiMo V2.6 Pro RL on 8× DGX Spark: 17.8 → 68.3 tok/s with DFlash
I’ve published our eight-Spark MiMo V2.6 Pro RL setup using official weights, vLLM and DFlash. Repository, setup and results Across four long structured-output tasks, throughput increased from 17.8 to 68.3 tokens/s, including prefill and request overhead. Both runs used the same patched runtime, w…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-23 07:18 · r/LocalLLM
Recipe & Patches: MiMo V2.6 Pro RL on 8× DGX Spark: 17.8 → 68.3 tok/s with DFlash