Qwen 3.8 27B on dual RTX 5060 Ti - 130 tok/s - full 256k context
I have been experimenting with Qwen 3.8 27B NVFP4 on my dual RTX 5060 Ti 16GB rig, for a total of 32GB VRAM. I'm running stock vLLM 0.30.0 with no patches, on headless Ubuntu, so no desktop is using any VRAM. This shows that it's possible to run Qwen 3.8 27B without an RTX 5090. My GPUs are power l…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-22 14:15 · r/LocalLLM
Qwen 3.8 27B on dual RTX 5060 Ti - 130 tok/s - full 256k context