vLLM Recipe for Qwen38 Flash Next NVFP4 TP=2 for RTX Pro 5000 72g
I couldn't find a recipe for this model on my hardware, so I used Hermes and Unsloth's 4-bit quant of the same model to cook up a vLLM recipe for the NVFP4 quant with PLE offloading. I've been running the model for about a week and it's taken everything I've thrown at it. Very happy with how it's p…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-25 22:53 · r/LocalLLaMA
vLLM Recipe for Qwen38 Flash Next NVFP4 TP=2 for RTX Pro 5000 72g