Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt
Real-world agent benchmark of AtomicChat/Qwen3.8-Flash-Next-GGUF , specifically the AD-3.84bpw-IQ4_XS-M64 quant, running on a small Ubuntu LLM server with 2× NVIDIA RTX 3060 12 GB . The goal was not maximum chat latency. The goal was to find out whether a very large MoE model could be useful as a q…
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-09-19 05:37 · r/LocalLLM
Has anyone used NVIDIA DGX Spark for serious cybersecurity workloads (Red Team, Blue Team, CTI, GRC)? - 2026-09-18 19:12 · r/LocalLLaMA
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark - 2026-09-18 13:16 · r/LocalLLM
Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt
More stories
- Microsoft director called AI scraping ‘the largest theft of labor in human history,’ while OpenAI head brands ChatGPT an ‘existential threat’ to publishers — revelations come from legal briefs filed in NYT lawsuit — Tom's Hardware
- The cloud outage that should terrify the CIO — InfoWorld AI
- Simulated students that make realistic mistakes help AI tutors learn faster — The Decoder
- If I buy the Pro version, will I automatically have access to GPT-6 Astra? — r/ChatGPTPro
- NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
- we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
- Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
- Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0) — r/huggingface
Get the daily brief of stories like this at 6:30 every morning →