llm performance community metric
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
my question about LLM performance We see a lot of posts about token prediction, token generation per second, etc. But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22–28 TPS, but I also see that the LLM does a lot of reasoning. A…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-13 12:52 · r/LocalLLaMA
llm performance community metric