Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator
Quick one for anyone wondering what a single 24 GB card is good for in an agent setup. I had Qwen 3.8 27B (Q4, llama.cpp) on one RTX 3090 do all the actual coding, and GPT-6.1 Sol in the cloud act as the orchestrator: it breaks the job into pieces, hands them out, and checks what comes back. The jo…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-30 00:42 · r/LocalLLM
Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator