Need advice 2x 3090 + 64GB DDR5 only getting 20t/s on Qwen3.8 Q4 llama.cpp
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
Recently swapped from Ollama to llama.cpp but haven't figured out how to run it efficiently. On Ollama I was getting 28t/s. Relatively new to this, advice welcome. I have also 4x more 3090s laying around. What's the best way to utilize them?
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-11 08:13 · r/LocalLLM
Need advice 2x 3090 + 64GB DDR5 only getting 20t/s on Qwen3.8 Q4 llama.cpp