Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up
This story is from 2026-10-09. It is preserved in the archive; the latest stories are on the live feed.
This article provides a step by step guide to serving every Gemma 4 size, E2B, E4B, 12B, 26B-A4B and 31B, on one AMD Instinct MI300X through vLLM in four weight formats, with each build timed across the same grid of request counts and prompt lengths on the same image. Every log, report and script i…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-09 13:46 · DEV Community — Machine Learning
Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up