When running Qwen3.8:27B is it better to have higher context or a secondary model for smaller tasks?
This story is from 2026-08-30. It is preserved in the archive; the latest stories are on the live feed.
I have 40GB of VRAM and 128GB of system memory. I am able to run Qwen3.8-27B Q4 with a 128k context length and I am running Qwen3-coder:30B for coding tasks. I’m wondering if I would do better to just give Qwen3.8 my full 40GB of VRAM for context or continue using it as an orchestrator for the codi…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-30 22:59 · r/LocalLLM
When running Qwen3.8:27B is it better to have higher context or a secondary model for smaller tasks?