Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers
We wanted to know if a local 27B can do the heavy lifting in an agent if something smarter does the planning. So we ran the same build three ways and wrote down everything. Hardware and setup: Qwen 3.8 27B UD-Q4_K_XL, single RTX 3090 24 GB, llama.cpp built with CUDA, 2 parallel slots, 64K context,…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-29 01:28 · r/LocalLLM
Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers