Help me choose hardware.
I need some advice. I want to choose hardware for local inference. Right now, I’m paying ~$30/day for rent on a Vast an RTX Pro 6000 SW 96GB, and I’m using Qwen3.8-Flash-Next-Q4_K_M.gguf at a speed of ~60 tok/s (I know this format isn’t efficient for this graphics card, but I’m limited in choosing…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-20 08:29 · r/LocalLLM
Help me choose hardware.