AINewsnow

What 2B tokens of coding-agent traffic taught us about using frontier and open models together

I run a small shared inference club serving Qwen 3.8 27B FP8 on one RTX PRO 6000 Blackwell. In the first week, it processed 2B tokens, almost entirely from coding agents working on real repositories. Two members used around 950M and 900M tokens each. One of my own tests was the browser port of Meda…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 13:58 · r/LocalLLM
    What 2B tokens of coding-agent traffic taught us about using frontier and open models together

More stories

  1. Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? — r/LocalLLaMA
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Layer Extract & Layer Remove Loras For Qwen Image 2.1 — r/StableDiffusion
  4. I built Slopus, a free, open-source desktop app for generating and editing AI videos locally (Minimax H3) — r/StableDiffusion
  5. Community reports say the first samples of Qwen 4 are already approaching Fable / Opus-level quality. — r/singularity
  6. Qwen-Image 2.1 Inpainting with LanPaint — alpha channel included — r/StableDiffusion
  7. A LoRA I made: AnyAngle LoRA for Qwen Image 2.1. Style-Aligned Arbitrary Camera Angles — r/StableDiffusion
  8. 2x Tesla P100, q6_k quant 50+tps. V2.0 — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →