Optimizing DGX Spark + RTX 4500 Pro with Qwen - slow TPS
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
I've got a DGX Spark and an RTX Pro 4500 in a whitebox AMD EPYC build (supermicro H11SSLi running proxmox, RTX Pro 4500 passed through to a docker host and shared between plex/frigate/llamacpp etc). ~6GB is actively consumed by non-LLM processes (frigate mostly) leaving 26GB of headroom. I'm workin…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-22 12:37 · r/LocalLLM
Optimizing DGX Spark + RTX 4500 Pro with Qwen - slow TPS