~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.
Stack Strata NVFP4 fork: github.com/sergqwer/strata-nvfp4 Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding W4A8 prefill on Blackwell Main engine flags: ./build/stra…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-05 11:09 · r/LocalLLM
~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.