AINewsnow

Strata: fetch experts before they're needed. Up to +34 % decode headroom measured on 2× 3090, more on smaller-VRAM cards

Quick background: I run Qwen3.8-Flash-Next (a big mixture-of-experts model) on 2x RTX 3090 with Strata ( https://github.com/Niko1221/Strata ). Like every MoE engine on consumer hardware, it can't fit all experts in VRAM. About half of them live in system RAM, and when a token needs one that isn't o…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-07 10:34 · r/LocalLLM
    Strata: fetch experts before they're needed. Up to +34 % decode headroom measured on 2× 3090, more on smaller-VRAM cards

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  3. Sharing AI progress in mathematics — OpenAI News
  4. Together Link: open models in the harness you already use. Start with one command today. — Together AI Blog
  5. OpenAI has dumped 722 maths papers – now it must clean up the mess — New Scientist Technology
  6. Google launches SynthID Detector, a website that lets users detect AI-generated image, video, and audio media across dozens of common file formats (Ivan Mehta/TechCrunch) — Techmeme
  7. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  8. Introducing the Decisions API — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →