Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
I wanted to share my successful setup for running a Qwen 3.8 27B model with a massive context window on a consumer 16GB GPU (RTX 4070 Ti SUPER). The goal was to fit everything into VRAM without sacrificing quality or speed. ๐ง Key Components Model: Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller from jrell on Hโฆ
Read the full story at r/LocalLLaMA โ
Timeline ยท 1 report
- 2026-08-29 12:50 ยท r/LocalLLaMA
Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)