Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this?
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
Hey r/LocalLLaMA , I wanted to share a milestone we just hit: we successfully deployed the Qwen 3.6 35B A3B model (NVFP4) on a single NVIDIA DGX Spark (GB10 / SM121) and managed to sustain some solid throughput under heavy load. I’m sharing the repo below, but I'm also hoping to get some feedback f…
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-09-19 05:37 · r/LocalLLM
Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)? - 2026-09-18 19:12 · r/LocalLLaMA
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark - 2026-09-17 17:47 · r/LocalLLM
Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this?