Sorting the vocabulary by spelling makes Gemma 3 270M 2.4× faster on CPU, and beats Gemma 4's own drafter token clustering
A two-level head (pick a group of tokens, then a token inside it) reads ~256× fewer output weights. The open question is how to group the tokens. Google's Gemma 4 drafters, FlashHead and others group by embedding similarity or by frequency. I tried the dumbest possible rule: sort the vocabulary by…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-09 22:25 · r/LocalLLM
Sorting the vocabulary by spelling makes Gemma 3 270M 2.4× faster on CPU, and beats Gemma 4's own drafter token clustering
More stories
- An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
- Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog
- Google Cloud introduces Gemini agent to change enterprise work — SiliconANGLE AI
- I need your help — r/learnmachinelearning
- Day 7 no Gemini 4 — r/GeminiAI
- Is Gemini Pro model down? — r/GeminiAI
- Rethinking access control for RAG with Amazon Quick and Amazon Bedrock — AWS Machine Learning Blog
- Thank you, Google — r/GeminiAI
Get the daily brief of stories like this at 6:30 every morning →