AINewsnow

GPU Poor Local usage 1 billion tokens a month?

just wondering how many tokens everyone with truly gpu poor setups are running. I started going all in on local since qwen 3.8 27b was launched and I am currently at 1.1 billion tokens in 3 weeks. Amazing results and truly seeing strides in intelligence. setup: 2 x 5060 ti 16gb vllm and 2 x 3060 12…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 02:28 · r/LocalLLM
    GPU Poor Local usage 1 billion tokens a month?

More stories

  1. Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
  2. Layer Extract & Layer Remove Loras For Qwen Image 2.1 — r/StableDiffusion
  3. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  4. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  5. I built Slopus, a free, open-source desktop app for generating and editing AI videos locally (Minimax H3) — r/StableDiffusion
  6. I tested Qwen Image 2.1, FLUX Klein 2, Krea 2, and Z Image in ComfyUI using the same prompts. Some interesting differences came out 👀 Full comparison here if you want to check it out! — r/comfyui
  7. Deepseek V4 Flash 0731 on m5 max 128gb — r/LocalLLM
  8. Community reports say the first samples of Qwen 4 are already approaching Fable / Opus-level quality. — r/singularity

Get the daily brief of stories like this at 6:30 every morning →