RANT: Model Cards Should List KV Cache Cost
I'm getting tired of looking at models -- especially small models -- that seem to fit on a modest GPU, only to find out that the model's attention mechanism is outdated and requires 64kB or more for each token of context. One of the key reasons that Ling-3.0-Tiny is a great model is that not only d…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-08 02:17 · r/LocalLLM
RANT: Model Cards Should List KV Cache Cost