Performance tune for gemma4-26b-a4b flash attention shape. by frobnitzem · Pull Request #28450 · ggml-org/llama.cpp · GitHub
Gemma 26B A4B speedup
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-22 17:39 · r/LocalLLaMA
Performance tune for gemma4-26b-a4b flash attention shape. by frobnitzem · Pull Request #28450 · ggml-org/llama.cpp · GitHub