AINewsnow

Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator

Quick one for anyone wondering what a single 24 GB card is good for in an agent setup. I had Qwen 3.8 27B (Q4, llama.cpp) on one RTX 3090 do all the actual coding, and GPT-6.1 Sol in the cloud act as the orchestrator: it breaks the job into pieces, hands them out, and checks what comes back. The jo…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 00:42 · r/LocalLLM
    Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator

More stories

  1. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  2. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  3. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  4. Qwen 3.8 is a workhorse — r/LocalLLaMA
  5. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy — r/LocalLLaMA
  7. Using GPT-6.1 Sol only as the planner and letting a local 27B write the code cut my API bill by 77% — r/ChatGPT
  8. Qwen Image 2.1 is so good and fun — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →