I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4
Hey guys, Last time I tested Qwen3.8-Flash-Next on its own. This time I put three Qwen3.8 checkpoints through the same 10 tests on the same RTX PRO 6000: RadixArk/Qwen3.8-27B-NVFP4 (dense) orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (dense, uncensored fine-tune) RadixArk/Qwen3.8-Flash-Next-NVFP4 (MoE)…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-02 21:47 · r/LocalLLaMA
I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4
More stories
- NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
- Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
- Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
- Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
- Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
- Introducing Clef: our open-source decision models, and new RL fine-tuning platform — Cloudflare Blog — AI
- Tavus unveils Griffin, the "first Human Interaction Model", which it says passed the "video Turing test", with 48% of users thinking it was human in live chats (@tavus) — Techmeme
- OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
Get the daily brief of stories like this at 6:30 every morning →