AINewsnow

Qwen3.8 27b - what is realistic tokens per second with 16 GB VRAM and 32 GB RAM?

I have NVidia RTX 4060 Ti with 16 GB VRAM and 32 GB of RAM. When I load Qwen 3.8 27b with llamacpp in 70k context size, I get speed around 6.5 tokens per second with ITQ4_XS quantz. On some threads people claim to get around 60-70 tokens per second with similar hardware, so I rather ask here what k…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-20 10:42 · r/LocalLLM
    Qwen3.8 27b - what is realistic tokens per second with 16 GB VRAM and 32 GB RAM?

More stories

  1. Alibaba Unveils AI Chip to Drive Global Data Center Buildout — Bloomberg AI
  2. SGLang Squeezes 78% More AI Throughput on NVIDIA's Blackwell GPUs — AlphaSignal
  3. Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters — The Decoder
  4. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  5. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  6. A Little Help — r/aivideo
  7. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  8. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial

Get the daily brief of stories like this at 6:30 every morning →