I published a 284x cost spread on one GPU. Re-ran it properly and it's 137x. Also found eager-mode numbers don't reproduce across machines.
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
Two weeks ago I posted that one model on one L4 showed a 284x range in cost per inference depending only on flags. I rebuilt the measurement this week and the honest number is 137x. Writing up what changed, because the reason it changed is more interesting than the number. WHAT I FIXED Prefix cachi…
Read the full story at r/LocalLLM ↗