Qwen3.8-Flash-Next NVFP4 at 256K context with Strata — 4,100+ prefill and up to 125 tok/s decode
I tested two Qwen3.8-Flash-Next NVFP4 checkpoints at a full 256K context using an experimental dual-GPU Strata configuration. Hardware: • RTX 4090 D 48 GB as the primary GPU • RTX 5070 Ti 16 GB as a helper GPU • Intel Core Ultra 7 265KF • 128 GB RAM • Linux • NVMe storage The RTX 4090 D handled the…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-03 02:54 · r/LocalLLM
Qwen3.8-Flash-Next NVFP4 at 256K context with Strata — 4,100+ prefill and up to 125 tok/s decode