New to local llms. Are these numbers normal?
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
I have a MacBook Air M5 with 16GB RAM, and I’m currently running unisloth Qwen 3.8 27B Q3_xxs on llama.cpp with KV cache 8. I’m pretty new to all this, so I’m wondering if ~9 t/s at 8k context sounds about right for this setup, and whether that’s enough for general chatting/inquiries. Also, should…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-07 04:45 · r/LocalLLM
New to local llms. Are these numbers normal?