Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
Bryan Shan / SemiAnalysis : Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Earn, Agen…
Read the full story at Techmeme ↗
Timeline · 1 report
- 2026-09-15 00:40 · Techmeme
Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)
More stories
- Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal
- DeepSeek’s Insane New Architecture — Two Minute Papers
- Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
- Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
- Ai used for chatbots — r/artificial
- I gave 6 different AIs the same 5 questions — r/AI_Agents
- PromptDeck v1.1.0 – open-source desktop app to benchmark local AND cloud LLMs side-by-side (Ollama, LM Studio + OpenRouter, Groq, DeepSeek…) — r/LocalLLaMA
- What’s your favorite AI model for coding right now? — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →