Gemma 4 E2B on an AMD MI300X: Which Weight Format Should You Serve?
This story is from 2026-10-08. It is preserved in the archive; the latest stories are on the live feed.
This article provides a step by step guide to serving ten weight formats of Gemma 4 E2B on one AMD Instinct MI300X through vLLM, with every build timed across a grid of request counts and prompt lengths on the same card, image and day. Every log, report and script is committed. On the MI300X, fp8 i…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-08 15:59 · DEV Community — Machine Learning
Gemma 4 E2B on an AMD MI300X: Which Weight Format Should You Serve?