Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1
Red Hat is proud to announce our results from the industry-standard MLPerf Inference v6.1 benchmark. This submission builds on our track record across recent rounds: In v5.1, we demonstrated cost-effective Llama-3.1-8B-FP8 inference with vLLM on NVIDIA H100 and L40S GPUs, and in v6.0 we delivered r…
Read the full story at Red Hat AI Blog ↗
Timeline · 1 report
- 2026-09-29 00:00 · Red Hat AI Blog
Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1