Finally got my 1x 5090 setup dialed in for agentic workflows: concurrent 920 t/s decode + 400 t/s prefill (qwen 3.8 27b)
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
Thought this might be the correct sub to share the rabbit hole I went into (and appreciate the numbers). Note: oneshotted the dashboard, it gets the job done. I have been optimizing throughput & quality now for a couple of weeks between longer runs to make this thing fly. Use case is mainly agentic…
Read the full story at r/LocalLLM ↗