vLLM Runs MiniMax M3 at 7.84x Faster on NVIDIA's Vera Rubin
vLLM lands day-0 support for NVIDIA Vera Rubin NVL72 with Rubin-tuned kernels and locality-aware MoE, hitting 7.8x the per-GPU throughput of GB200.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-10-10 01:49 · AlphaSignal
vLLM Runs MiniMax M3 at 7.84x Faster on NVIDIA's Vera Rubin