A very confusing report from Puget Systems
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Just to name a few: running Qwen3 8B on a 32GB GPU running Qwen3.6-27B Q4_K_M on 2 x R9700 quote: "each prompt was sized at 500 input and 500 output tokens" for a full system that costs $18,775?? I don't understand what they are doing. Am I reading something wrong?
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-01 03:19 · r/LocalLLaMA
A very confusing report from Puget Systems