Our self-hosted inference cost per token only beat the API once the GPUs had night work.
Our API bill for two internal tools had crept up to around 9k a month, so in spring we bought a 4 card box, put a mid size Qwen on vLLM and moved everything over. The spreadsheet said our cost per token would be about a fifth of the API which is what got finance to sign off. Three months later I pu…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-07 21:13 · r/LocalLLM
Our self-hosted inference cost per token only beat the API once the GPUs had night work.