AINewsnow

CoreWeave Delivers Breakthrough AI Performance with NVIDIA GB200 and H200 GPUs in MLPerf Inference v5.0

This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.

CoreWeave achieved top MLPerf v5.0 AI inference results, delivering 800 TPS on Llama 3.1 405B with NVIDIA GB200 and 33,000 TPS on Llama 2 70B with H200 GPUs, marking significant performance gains.

Read the full story at CoreWeave Blog ↗

Timeline · 1 report

  1. 2026-09-08 13:59 · CoreWeave Blog
    CoreWeave Delivers Breakthrough AI Performance with NVIDIA GB200 and H200 GPUs in MLPerf Inference v5.0

More stories

  1. Which models you run on your Nvidia v100? — r/LocalLLM
  2. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  3. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  4. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  5. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  6. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  7. focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737) — r/LocalLLaMA
  8. I ran Opencode and PI against the same local model on 3 identical projects, same prompts, same hardware... — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →