Over 1000 tok/s decode for 27B on my 4x MI100 server
So here is my server with four MI100, each with 32 GB of VRAM, with their infinity fabric bridge directly connecting all of them. MI100 is not well supported out of box in most software, so I made a fork of vLLM to get the performance I was hoping for. VLLM was running at 15 tok/s stock on this mac…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-28 21:32 · r/LocalLLM
Over 1000 tok/s decode for 27B on my 4x MI100 server