Inception's Mercury 2.5 Hits 1,107 Tokens per Second, Beating Autoregressive Models
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
Inception's new diffusion-based LLM hits 1,107 tokens per second on NVIDIA GPUs while boosting quality 40% over Mercury 2.
Read the full story at AlphaSignal ↗
Timeline · 2 reports
- 2026-09-09 11:48 · TestingCatalog AI News
Inception launches Mercury 2.5 at 1,107 tokens per second - 2026-09-08 16:45 · AlphaSignal
Inception's Mercury 2.5 Hits 1,107 Tokens per Second, Beating Autoregressive Models