Cohere's Open-Source Megakernel Beats vLLM by 1.58x on H100
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
Cohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-09-08 19:44 · AlphaSignal
Cohere's Open-Source Megakernel Beats vLLM by 1.58x on H100