AINewsnow

Are these the best llama.cpp settings for Qwen 3.8 on a 24 GB RTX 4090?

This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.

I’m looking for feedback from people familiar with Qwen 3.8 and llama.cpp. Are these sensible settings, or are there better choices for quality, speed, VRAM usage, and long-context performance? My use is coding and recurring/scheduled agentic tasks. Hardware and server GPU: NVIDIA RTX 4090 24 GB Ba…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-08-20 20:53 · r/LocalLLM
    Are these the best llama.cpp settings for Qwen 3.8 on a 24 GB RTX 4090?

More stories

  1. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  2. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  3. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  4. My Version of Jev running locally, playing doom. — r/LocalLLM
  5. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  6. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  7. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  8. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →