AINewsnow

Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1

Red Hat is proud to announce our results from the industry-standard MLPerf Inference v6.1 benchmark. This submission builds on our track record across recent rounds: In v5.1, we demonstrated cost-effective Llama-3.1-8B-FP8 inference with vLLM on NVIDIA H100 and L40S GPUs, and in v6.0 we delivered r…

Read the full story at Red Hat AI Blog ↗

Timeline · 1 report

  1. 2026-09-29 00:00 · Red Hat AI Blog
    Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  4. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  5. AI leaders talk latest models, tech risks at Trump lunch — Semafor Technology
  6. Trump and Johnson to meet with tech CEOs on AI risk — Politico Technology
  7. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  8. Inference Engines will become a series of one-offs — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →