Qwen3.8-Flash-Next, dual RTX 3090s, NVLink: ~3,200 t/s prefill, ~115 t/s decode
Background: Strata ( https://github.com/Niko1221/Strata ) runs the Qwen3.8-Flash-Next GGUF very well, but on one card. The built-in multi-GPU mode splits layers, which didn't do much for decode on a 3090 pair. My engine here is Strata v0.1.38 plus 29 commits that instead keep the whole model on the…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-06 04:42 · r/LocalLLM
Qwen3.8-Flash-Next, dual RTX 3090s, NVLink: ~3,200 t/s prefill, ~115 t/s decode
More stories
- Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
- Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
- Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen — r/LocalLLaMA
- OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
- Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
- Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
- Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
- can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →