Are these the best llama.cpp settings for Qwen 3.8 on a 24 GB RTX 4090?
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
I’m looking for feedback from people familiar with Qwen 3.8 and llama.cpp. Are these sensible settings, or are there better choices for quality, speed, VRAM usage, and long-context performance? My use is coding and recurring/scheduled agentic tasks. Hardware and server GPU: NVIDIA RTX 4090 24 GB Ba…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 20:53 · r/LocalLLM
Are these the best llama.cpp settings for Qwen 3.8 on a 24 GB RTX 4090?