AINewsnow

Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt

Real-world agent benchmark of AtomicChat/Qwen3.8-Flash-Next-GGUF , specifically the AD-3.84bpw-IQ4_XS-M64 quant, running on a small Ubuntu LLM server with 2× NVIDIA RTX 3060 12 GB . The goal was not maximum chat latency. The goal was to find out whether a very large MoE model could be useful as a q…

Read the full story at r/LocalLLM ↗

Timeline · 3 reports

  1. 2026-09-19 05:37 · r/LocalLLM
    Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)?
  2. 2026-09-18 19:12 · r/LocalLLaMA
    Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark
  3. 2026-09-18 13:16 · r/LocalLLM
    Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt

More stories

  1. Microsoft director called AI scraping ‘the largest theft of labor in human history,’ while OpenAI head brands ChatGPT an ‘existential threat’ to publishers — revelations come from legal briefs filed in NYT lawsuit — Tom's Hardware
  2. The cloud outage that should terrify the CIO — InfoWorld AI
  3. Simulated students that make realistic mistakes help AI tutors learn faster — The Decoder
  4. If I buy the Pro version, will I automatically have access to GPT-6 Astra? — r/ChatGPTPro
  5. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  6. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  7. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  8. Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0) — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →