Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark
Planning, coding, testing = 8h total. Stack: VSCode Copilot in autopilot mode + SGLang Stats: ∼10k lines generated, ∼800k tokens consumed Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-19 05:37 · r/LocalLLM
Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)? - 2026-09-18 19:12 · r/LocalLLaMA
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark