KV Cache Math: Why Llama 3.1 8B at 128K Context Won't Fit in 24GB
I quantized Llama 3.1 8B down to Q4 so it would fit comfortably on my 24GB card. The weights came out under 5 GB. Then I set the context to 128K, because the model card says it supports 128K, and the model promptly spilled onto the CPU and generated text at the speed of a fax machine. The weights w…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-27 05:35 · DEV Community — Machine Learning
KV Cache Math: Why Llama 3.1 8B at 128K Context Won't Fit in 24GB