Concurrent requests on M5 Ultra?
Has anyone experimented with concurrent requests on their M5 Ultra with ds4 or omlx? I’m curious how the prefill/token gen speeds hold up with multiple concurrent requests. I’m within the return window for my 2x DGX sparks but i really like how concurrency feels on them. I’ve been running 4 Qwen 3.…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-02 16:51 · r/LocalLLM
Concurrent requests on M5 Ultra?