Improve token per second without touching quant
Spent the past month tweaking and experimenting with many different numbers to achieve 30tps. Hardware: -Rx6700xt 12gb vram (AMD) -2x16 ddr4 3200 ram -r5 5600x -llama.cpp vulkan sdk -window11 (no wsl switching since im not used to the environment) Im looking for any improvements to achieve maybe 40…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-10 12:05 · r/LocalLLaMA
Improve token per second without touching quant