AINewsnow

Over 1000 tok/s decode for 27B on my 4x MI100 server

So here is my server with four MI100, each with 32 GB of VRAM, with their infinity fabric bridge directly connecting all of them. MI100 is not well supported out of box in most software, so I made a fork of vLLM to get the performance I was hoping for. VLLM was running at 15 tok/s stock on this mac…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-28 21:32 · r/LocalLLM
    Over 1000 tok/s decode for 27B on my 4x MI100 server

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. OpenAI Scraps Debut of Latest Astra Model Over Safety Risks — Bloomberg AI
  4. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  5. Meta launches enterprise AI business seeking to cash in on vast spending — Financial Times AI
  6. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  7. AMD to Buy Fei-Fei Li’s World Labs AI Startup for $8.2 Billion — Bloomberg AI
  8. Scoop: Anthropic's Dario Amodei to have White House dinner with Trump — Axios AI+

Get the daily brief of stories like this at 6:30 every morning →