AINewsnow

Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up

This story is from 2026-10-09. It is preserved in the archive; the latest stories are on the live feed.

This article provides a step by step guide to serving every Gemma 4 size, E2B, E4B, 12B, 26B-A4B and 31B, on one AMD Instinct MI300X through vLLM in four weight formats, with each build timed across the same grid of request counts and prompt lengths on the same image. Every log, report and script i…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-09 13:46 · DEV Community — Machine Learning
    Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  3. OpenAI Decisions API now available on AI Gateway — Vercel Blog
  4. Introducing Playground: Create and play custom games — Google AI Blog
  5. Anthropic bans ‘abusive or cruel behavior’ toward Claude — The Verge AI
  6. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  7. Fired OpenAI safety researchers dispute their dismissals in open letter — Engadget
  8. Sophos cuts threat investigation time by 96% with OpenAI Daybreak — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →