what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?
we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a "twin 3090 and 128GB RAM" recipe around here recently that could do tensor parallel and I can't find it anymore.
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-19 05:37 · r/LocalLLM
Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)? - 2026-09-18 19:12 · r/LocalLLaMA
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark - 2026-09-18 13:16 · r/LocalLLM
Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt - 2026-09-18 01:36 · r/LocalLLaMA
what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?