AINewsnow

2B tokens of local Qwen3.8 as Claude Code's worker: 72% less spend on claude tokens, what a 27B NVFP4 carries, and the tool that makes it all seamless

This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.

On a real job with hidden checks, letting the expensive model plan and review while a local Qwen typed, cost 72% less and used 85% fewer expensive-model tokens. Day to day, every expensive token the head spends buys 20 to 50 tokens of cheap local work, and the tighter that ratio, the more efficient…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-12 22:30 · r/LocalLLM
    2B tokens of local Qwen3.8 as Claude Code's worker: 72% less spend on claude tokens, what a 27B NVFP4 carries, and the tool that makes it all seamless

More stories

  1. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  2. Ho creato un piccolo benchmark "test nascosti + revisione del codice" e l'ho eseguito su Claude Sonnet 5, Claude Opus 4.6 e una versione locale di Qwen 3.8 27B (Unsloth Q6 - Qwen3.8-27B-UD-Q6_K.gguf). Risultati + cosa li ha effettivamente differenziati — r/LocalLLM
  3. Agentic Orchestration with Local and Cloud Models — r/LocalLLM
  4. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/ArtificialInteligence
  5. Best workflow to orchestrate local LLMs (Qwen 27B) + Cloud subscriptions (Claude/Codex) without burning tokens? — r/LocalLLM
  6. Qwen Next 3.8 / Claude Opus level local model - What to buy in order to deploy? — r/LocalLLaMA
  7. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  8. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →