How are you measuring model cost when the task needs retries, tools, and human review?
I have been trying to move model comparisons away from token price alone and toward the cost of an accepted work product. One current example is Fireworks' reported Ember-1 comparison: total tokens moved from 49.3K to 29.9K, while the task score moved from 0.751 to 0.753. That is a promising token-…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-03 04:56 · r/LocalLLM
How are you measuring model cost when the task needs retries, tools, and human review?