AINewsnow

Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm

I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and configurations on e…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-22 09:52 · r/LocalLLaMA
    Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. Transformers now runs llama.cpp quants — Hugging Face Blog
  3. Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken — r/LocalLLaMA
  4. I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
  5. Is llama.cpp meant to be slow at long context, even when you aren't using that context? — r/LocalLLaMA
  6. Offline Ghostwriter Studio – A 100% private, offline IDE for novelists powered by local llama-server (No subscription, 133 MB installer) — r/LocalLLM
  7. Looking for LM Studio replacement, tired of the nonsense. — r/LocalLLM
  8. PXA v2026.09.20 — my inference engine for old Teslas (P100 / V100 / 1080 Ti): Gemma 4 MoE, tensor split on by default, and ahead of stock llama.cpp on every cell on my rig — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →