Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers)
I thought I'd share my daily-driver setup for Qwen3.8-27B on two 3090s. Most dual-3090 numbers I see are for 4-bit models, but this one keeps 8-bit weights (INT8 W8A16, Q8_0-class fidelity). It still decodes about 2x faster than the llama.cpp Q8_0 + MTP setup it replaced on the same box. Hopefully…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-30 06:37 · r/LocalLLM
Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers)
More stories
- Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
- Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
- Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
- Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
- Final-year student in India trying to break into generative-model inference optimization — roadmap feedback? — r/MLQuestions
- Model Registry (RTX 4090) — r/LocalLLM
- dual 20gb 3080 and 4070ti local llm — r/LocalLLM
- Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →