AINewsnow

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware get…

Read the full story at NVIDIA Blog ↗

Timeline · 2 reports

  1. 2026-09-16 15:26 · CoreWeave Blog
    CoreWeave Leads Cloud Providers in MLPerf® Inference v6.1 Performance with NVIDIA Blackwell Ultra
  2. 2026-09-16 15:00 · NVIDIA Blog
    NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

More stories

  1. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  2. Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters — FP4 performance of the Ascend 960PR doubles expectations — Tom's Hardware
  3. Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. — r/LocalLLaMA
  4. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  5. Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0) — r/huggingface
  6. Zoom’s CEO agrees with Bill Gates, Jensen Huang, and Jamie Dimon: A 3-day workweek is coming soon thanks to AI — Fortune AI
  7. Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. — r/LocalLLaMA
  8. FREE AI TRAINING CREDIT — r/learnmachinelearning

Get the daily brief of stories like this at 6:30 every morning →