[R] Fathom: letting each query choose how many bits of each key channel to read, for sparse decoding over an offloaded KV cache
Sparse attention fetches the top-k keys, but to know which k you first have to rank all n. Every method does that by scanning a cheap compressed copy of every key, and every one of them fixes the bit depth of that copy at design time: SparQ reads 16 or 32 channels at full depth, Double Sparsity rea…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-29 14:25 · r/learnmachinelearning
[R] Fathom: letting each query choose how many bits of each key channel to read, for sparse decoding over an offloaded KV cache
More stories
- NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
- How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
- OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
- Anthropic warns of ‘existential risks to humanity’ in IPO prospectus — Financial Times AI
- Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
- OpenAI launches Dots, always-on agents powered by GPT-6 Astra with their own cloud computer, in ChatGPT for Pro, Business Premium, and Enterprise users (Rachel Metz/Bloomberg) — Techmeme
- OpenAI Scraps Release of New AI Model Over Safety Concerns — Wall Street Journal Technology
- OpenAI DevDay 2026: The biggest news and announcements — The Verge AI
Get the daily brief of stories like this at 6:30 every morning →