AINewsnow

Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this?

This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.

Hey r/LocalLLaMA , I wanted to share a milestone we just hit: we successfully deployed the Qwen 3.6 35B A3B model (NVFP4) on a single NVIDIA DGX Spark (GB10 / SM121) and managed to sustain some solid throughput under heavy load. I’m sharing the repo below, but I'm also hoping to get some feedback f…

Read the full story at r/LocalLLM ↗

Timeline · 3 reports

  1. 2026-09-19 05:37 · r/LocalLLM
    Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)?
  2. 2026-09-18 19:12 · r/LocalLLaMA
    Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark
  3. 2026-09-17 17:47 · r/LocalLLM
    Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this?

More stories

  1. what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn? — r/LocalLLaMA
  2. Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. — r/LocalLLaMA
  3. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  4. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  5. Microsoft exec called AI scraping the “largest theft of labor in human history” — Ars Technica AI
  6. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  7. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  8. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →