AINewsnow

What 125B parameters at 100 tok/s on one GPU costs to reproduce

Building AI that practitioners can actually run The pipeline of open-source work this month points to a single theme: practitioners are less interested in abstract capability claims than in systems they can reproduce, govern, and maintain. The Strata project shows Qwen 3.8 Flash Next (125B) running…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-05 01:53 · DEV Community — Machine Learning
    What 125B parameters at 100 tok/s on one GPU costs to reproduce

More stories

  1. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  2. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  3. One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui
  4. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  5. My frontier class agent fact-checks my local AI before I grade it. How do you grade your Agents and LLMs? — r/AI_Agents
  6. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  7. ComfyUI Qwen image 2.1 Enhancer (Two nodes) — r/StableDiffusion
  8. I built Ninfer 4080 for 16GB class GPUs — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →