Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16
This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.
This article provides a step by step guide to shrinking Google's quantization-aware-trained (QAT) Gemma 4 E2B to 4-bit weights end to end, embedding tables included, and serving it with vLLM on one Tesla T4 attached to a Compute Engine VM. It compares the result with the bf16 reference and with Goo…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-30 13:00 · DEV Community — Machine Learning
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16