1,107 Tokens Per Second: The LLM That Doesn't Type
On September 8, 2026, Inception Labs announced Mercury 2.5 — which the company describes as the largest diffusion language model ever trained. The headline number: 1,107 tokens per second on widely available NVIDIA GPUs, at quality the company says matches the cost-optimized frontier tier (GPT-5.6…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-05 21:54 · DEV Community — Machine Learning
1,107 Tokens Per Second: The LLM That Doesn't Type