[Release] Breaking the KV-Cache Memory Wall: We Compressed ISOM-R1 Down to 285 Lines of Vectorized PyTorch (Open Weights for 40B)
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
TL;DR : At long context (64K–128K+), KV-cache VRAM surpasses model weights, making local inference unaffordable on consumer hardware. We are open-sourcing the ISOM-R1 family (Isometric State-Space & Orthogonal Manifolds). It bounds KV-cache memory into a fixed INT8 budget via quantum Gram-matrix ro…
Read the full story at r/huggingface ↗