AINewsnow

Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s

This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.

My setup: RTX 5070Ti, 16GB DDR5, Ryzen 7 9700X, Windows 11 The model: Unsloth’s UD-IQ4_XS https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/main/Qwen3.8-27B-UD-IQ4_XS.gguf Qwen3.8 reasons exclusively in English unless the system prompt explicitly directs it, and my workload doesn’t have the mode…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-09 23:42 · r/LocalLLaMA
    Running qwen 3.8 27B iq3 xxs on RTX 3060.
  2. 2026-09-08 15:53 · r/LocalLLM
    Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s

More stories

  1. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  2. Qwen/Qwen-Image-2.1 · Hugging Face — r/StableDiffusion
  3. Qwen Image 2.1 Releasing Tomorrow — r/StableDiffusion
  4. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  5. Alibaba releases Qwen-Image-2.1, a 7B open-weight model it says outperforms most closed-source models, with native transparency and up to ten reference images (Qwen) — Techmeme
  6. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  7. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  8. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial

Get the daily brief of stories like this at 6:30 every morning →