AINewsnow

what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?

we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a "twin 3090 and 128GB RAM" recipe around here recently that could do tensor parallel and I can't find it anymore.

Read the full story at r/LocalLLaMA ↗

Timeline · 4 reports

  1. 2026-09-19 05:37 · r/LocalLLM
    Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)?
  2. 2026-09-18 19:12 · r/LocalLLaMA
    Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark
  3. 2026-09-18 13:16 · r/LocalLLM
    Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt
  4. 2026-09-18 01:36 · r/LocalLLaMA
    what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?

More stories

  1. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  2. The cloud outage that should terrify the CIO — InfoWorld AI
  3. Simulated students that make realistic mistakes help AI tutors learn faster — The Decoder
  4. If I buy the Pro version, will I automatically have access to GPT-6 Astra? — r/ChatGPTPro
  5. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  6. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  7. AI skills — r/AI_Agents
  8. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →