Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and configurations on e…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-22 09:52 · r/LocalLLaMA
Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
More stories
- Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
- Transformers now runs llama.cpp quants — Hugging Face Blog
- Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken — r/LocalLLaMA
- I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
- Is llama.cpp meant to be slow at long context, even when you aren't using that context? — r/LocalLLaMA
- Offline Ghostwriter Studio – A 100% private, offline IDE for novelists powered by local llama-server (No subscription, 133 MB installer) — r/LocalLLM
- Looking for LM Studio replacement, tired of the nonsense. — r/LocalLLM
- PXA v2026.09.20 — my inference engine for old Teslas (P100 / V100 / 1080 Ti): Gemma 4 MoE, tensor split on by default, and ahead of stock llama.cpp on every cell on my rig — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →