Strata: Qwen 3.8 Flash next in loop
Sto eseguendo Qwen 3.8 Flash Next IQ3_XXS su Radeon W7800 48gb Vram e 64gb Ram ddr5. Ho scaricato tutto da github e lanciato l'eseguibile. Ho scelto solo la quantizzazione, il contesto 128k e la kv 8q. Primo compito per testarlo: fare un audit di un file js di 27kb e alla fine fare un report comple…
Read the full story at r/LocalLLM ↗
Timeline · 6 reports
- 2026-10-07 22:26 · r/LocalLLM
Is anyone running Qwen Flash Next on Strata with q6 or higher? - 2026-10-07 22:02 · r/LocalLLM
Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next - 2026-10-07 20:43 · r/LocalLLaMA
Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B? - 2026-10-06 14:58 · r/LocalLLaMA
Qwen3.8-Flash-Next on Strata - 2026-10-06 10:37 · r/LocalLLaMA
NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s - 2026-10-05 17:38 · r/LocalLLM
Strata: Qwen 3.8 Flash next in loop
More stories
- Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA
- RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
- Uncensored models: what they are and how they work, with a file-by-file hash check of one uncensored Qwen3.8-Flash-Next upload against Qwen's original — r/LocalLLM
- Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM
- Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
- A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
- Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2). One-command install, everything pinned. — r/LocalLLM
- Qwen 2.1 nails product references but struggles with people – combine with Krea 2 for realistic lifestyle scenes? — r/comfyui
Get the daily brief of stories like this at 6:30 every morning →