AINewsnow

KV cache quantization test: up to 2.6 GB less kv_cache, no quality loss I could measure (Qwen3.8-27B, RX 7900 XTX)

I wanted to know what KV cache quantization actually costs on my 24 GB card. So I ran Qwen3.8-27B (UD-Q4_K_M) at 65k and some other context sizes on my RX 7900 XTX with f16, q8_0, q5_0 and q4_0 KV cache. llama.cpp via Unsloth Studio. What I found: Memory: the f16 cache is 5.9 GB at 65k, and q4_0 ha…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-25 19:36 · r/LocalLLM
    KV cache quantization test: up to 2.6 GB less kv_cache, no quality loss I could measure (Qwen3.8-27B, RX 7900 XTX)

More stories

  1. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  2. model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp — r/LocalLLaMA
  3. My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context — r/LocalLLaMA
  4. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  5. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  6. Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM
  7. Qwen3.8 27B on Intel X7 358h + B390, with pi + llama.cpp surprised by its own RAM speed — r/LocalLLM
  8. 2x Radeon AI PRO R9700 + Qwen3.8-27B: 31 t/s on Windows → 113 t/s on Linux/vLLM. The fix was an M.2 riser, because the chipset slot doesn't do PCIe atomics. — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →