Qwen 3.8 27B FP8 on one RTX PRO 6000: 2B tokens of shared coding-agent traffic
I've been running Qwen 3.8 27B FP8 on vLLM (prefix caching + speculative decoding) on a single RTX PRO 6000 Blackwell 96GB, shared by a small group of developers who mostly use it for coding agents. In its first week it processed 2B tokens. Two users alone went through ~950M and ~900M. What the wor…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-23 12:51 · r/LocalLLM
Qwen 3.8 27B FP8 on one RTX PRO 6000: 2B tokens of shared coding-agent traffic