NVIDIA Groq 3 LPX Hits 3,400 Tokens/s on Gemma 4 31B
This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.
NVIDIA's Groq 3 LPX claims 3,400 tokens/s on Gemma 4 31B, with Nebius first to deploy. The 4x responsiveness claim targets agent latency. NVIDIA's Groq 3 LPX hit 3,400 output tokens/s on Gemma 4 31B, per a tweet from @kimmonismus. The dedicated token-generation accelerator is now in full production…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-08-24 22:26 · DEV Community — AI
NVIDIA Groq 3 LPX Hits 3,400 Tokens/s on Gemma 4 31B