Slow and Smart or Fast and Decent
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
So, with an agentic harness like Hermes Agent, the system can retry a task until it succeeds, which makes model speed/latency pretty important, especially when Hardware is the bottleneck. Would it therefore be more beneficial to use a smart, capable model that can run at, say, maybe 150 tokens/sec,…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-05 23:29 · r/LocalLLM
Slow and Smart or Fast and Decent