ConvRot Quant method now in llama-cpp-turboquant
This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.
It started here , and now https://github.com/TheTom/llama-cpp-turboquant/ has it. Imagine a Q6 quant with nearly Q8 KLD/PPL. Q6_CR and Q5_CR have a slight improvement over their base counterparts. Also while you are there check out --moe-cache auto to help improve running MoE models bigger than you…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-24 01:16 · r/LocalLLaMA
ConvRot Quant method now in llama-cpp-turboquant