AINewsnow

Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers

We wanted to know if a local 27B can do the heavy lifting in an agent if something smarter does the planning. So we ran the same build three ways and wrote down everything. Hardware and setup: Qwen 3.8 27B UD-Q4_K_XL, single RTX 3090 24 GB, llama.cpp built with CUDA, 2 parallel slots, 64K context,…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 01:28 · r/LocalLLM
    Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers

More stories

  1. PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill — r/LocalLLM
  2. Qwen 3.8 is a workhorse — r/LocalLLaMA
  3. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy — r/LocalLLaMA
  5. I built an open-source Prompt Engine & Screenplay Studio to solve video diffusion token drops & character drift (MiniMax H3 / Kling / Maestro) — r/PromptEngineering
  6. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  7. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  8. Qwen3.8 27B on Intel X7 358h + B390, with pi + llama.cpp surprised by its own RAM speed — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →