New beellama fork 76% faster tg with kvarn KV quants
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
When using kvarn quants at low context depth tg speed is similar to llama.cpp on and equivalent qx_x quant. However, as context depth grows kvarn tanks your tg speed. This fork optimises kvarn to have similar or better performance ay high context depths than llama.cpp at an equivalent qx_x quant an…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-31 20:03 · r/LocalLLaMA
New beellama fork 76% faster tg with kvarn KV quants