DeepSeek MLA: 70 GB of KV Cache at 1M Tokens
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
DeepSeek-V3's published config.json caches 576 numbers per token, per layer. Across its 61 layers in bf16, that is 70,272 bytes — about 70 KB of KV cache for every token sitting in the window. Fill a million-token context with a single sequence and you are holding roughly 70 GB of cache, before mod…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-07 04:43 · DEV Community — AI
DeepSeek MLA: 70 GB of KV Cache at 1M Tokens