2× RTX PRO 6000 Blackwell Server Edition: Qwen3.8-Flash-Next, two TP1 replicas, 1M context, FP8 KV + shared Mooncake RAM cache
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
We’re running Qwen3.8-Flash-Next NVFP4 on two RTX PRO 6000 Blackwell Server Edition GPUs, with a separate model instance on each card. Each instance is TP1. A customized gateway routes requests between the two replicas using least-inflight routing, and both instances can reuse prefix caches through…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-15 22:04 · r/LocalLLM
RTX PRO 6000 Blackwell hitting 85–92°C under sustained AI workloads. Is this normal? - 2026-09-15 12:46 · r/LocalLLM
2× RTX PRO 6000 Blackwell Server Edition: Qwen3.8-Flash-Next, two TP1 replicas, 1M context, FP8 KV + shared Mooncake RAM cache