49 tok/s 4060ti
Getting 40-50 toks/s qwen3.8 unsloth iq4 xs with ollama off of a 4060ti 16gb. Context is 32k but it doesn’t seem an issue with autocompact in vscode. Using claude to manage and direct qwen. Code reviews are clean and it’s fast enough for my workflow. Qwen is saving me money as I have gone from $300…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-25 01:43 · r/LocalLLM
49 tok/s 4060ti