AINewsnow

Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure

I put an old RTX 3090 to work as a local coding agent with Qwen3.8-27B, llama.cpp, and OpenCode. Here are the setup and results, including what failed. This is a summary of my own blog post, linked below. Setup RTX 3090 24GB, Ryzen 7 5800X, 64GB RAM, Ubuntu. Qwen3.8-27B Q4_K_M weights (~16.8GB), al…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-01 20:04 · r/LocalLLaMA
    Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure

More stories

  1. RTX 5070 TI and 32GB RAM DDR4, WHAT IS THE BEST VIDEO MODEL I COULD USE? — r/comfyui
  2. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant — r/LocalLLaMA
  4. From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck — r/LocalLLaMA
  5. Need maybe say "Use llama.cpp" — r/LocalLLaMA
  6. I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra) — r/LocalLLM
  7. Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality — r/LocalLLM
  8. The Rise of Overfit Inference Engines — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →