Qwen3.8 27b - what is realistic tokens per second with 16 GB VRAM and 32 GB RAM?
I have NVidia RTX 4060 Ti with 16 GB VRAM and 32 GB of RAM. When I load Qwen 3.8 27b with llamacpp in 70k context size, I get speed around 6.5 tokens per second with ITQ4_XS quantz. On some threads people claim to get around 60-70 tokens per second with similar hardware, so I rather ask here what k…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-20 10:42 · r/LocalLLM
Qwen3.8 27b - what is realistic tokens per second with 16 GB VRAM and 32 GB RAM?