AINewsnow

200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)

Strata runs Qwen3.8-Flash-Next on gaming PCs, but with 32 GB of RAM its low-RAM mode only works on one GPU. Split the model across two cards and the experts that don't fit in VRAM get read from the SSD. My 4060 Ti sat next to the 5080 doing nothing. So I forked it. The RAM copy of the experts now w…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-03 14:50 · r/LocalLLM
    200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)

More stories

  1. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  2. OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
  3. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  4. Introducing Oscilloscope Diffusion — r/comfyui
  5. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  6. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  7. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost
  8. Everything we launched during Birthday Week 2026 — Cloudflare Blog — AI

Get the daily brief of stories like this at 6:30 every morning →