AINewsnow

People with RTX PRO 6000, what tokens per second are you getting with Qwen 3.8 27B?

This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.

I have a dual RTX PRO 4000 setup. I get around 35 tokens per second with Qwen 3.8 27B Q6. But above 100k context, it drops down to around 20. I was considering an upgrade in the near future and I’m just curious what numbers people with the RTX 6000 are getting. On paper the RTX 6000 is paper becaus…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-01 18:14 · r/LocalLLM
    People with RTX PRO 6000, what tokens per second are you getting with Qwen 3.8 27B?

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  3. Alibaba releases Qwen-Image-2.1, a 7B open-weight model it says outperforms most closed-source models, with native transparency and up to ten reference images (Qwen) — Techmeme
  4. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  5. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  6. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial
  7. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  8. M2 Mac ultra128gb Qwen flash next — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →