Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't
This article provides a step by step guide to repacking Google's quantization-aware-trained (QAT) Gemma 4 weights for vLLM and serving them on one Google Cloud TPU v5e chip, with every build scored for classification, math, tool calling, throughput and long prompts. Every per-record output, log and…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 2 reports
- 2026-10-02 17:59 · DEV Community — Machine Learning
Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second - 2026-10-02 16:59 · DEV Community — Machine Learning
Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't