AINewsnow

A 125B Model at 100 tok/s on One RTX 4090? Here's What the HN Hype Leaves Out

This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.

A repo called Strata just hit 788 points on Hacker News with a headline that sounds like a typo: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100 tok/s. A 125B-parameter model. On a gaming GPU. At 100 tokens per second. Is it real? Yes. Is it what the title implies? Not quite.…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-05 09:02 · DEV Community — AI
    A 125B Model at 100 tok/s on One RTX 4090? Here's What the HN Hype Leaves Out

More stories

  1. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  2. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost
  3. One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui
  4. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  5. My frontier class agent fact-checks my local AI before I grade it. How do you grade your Agents and LLMs? — r/AI_Agents
  6. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  7. ComfyUI Qwen image 2.1 Enhancer (Two nodes) — r/StableDiffusion
  8. I built Ninfer 4080 for 16GB class GPUs — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →