What 2B tokens of coding-agent traffic taught us about using frontier and open models together
I run a small shared inference club serving Qwen 3.8 27B FP8 on one RTX PRO 6000 Blackwell. In the first week, it processed 2B tokens, almost entirely from coding agents working on real repositories. Two members used around 950M and 900M tokens each. One of my own tests was the browser port of Meda…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-29 13:58 · r/LocalLLM
What 2B tokens of coding-agent traffic taught us about using frontier and open models together