Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second
This article provides a step by step guide to repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for vLLM and serving them on one Google Cloud TPU v5e chip, with every build scored for classification, math, tool calling, throughput and long prompts. Every per-record output, log and…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-02 17:59 · DEV Community — Machine Learning
Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second