AINewsnow

Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram

mrdefaultuser@team-green : ~ $ /srv/ai/llama.cpp/build/bin/llama-server \ -m /srv/ai/models/qwen/Qwen3.8-Flash-Next-Q8_0/Qwen3.8-Flash-Next-Q8_0-00001-of-00002.gguf \ -c 262144 \ -np 1 \ -ngl 64 \ --cpu-moe \ --lazy-mode on \ --load-mode auto \ --agent \ --tools all \ --host 0.0.0.0 \ --port 8080 E…

Read the full story at r/LocalLLM ↗

Timeline · 5 reports

  1. 2026-10-06 05:39 · r/LocalLLaMA
    unsloth/Qwen3.8-Flash-Next-GGUF is being updated
  2. 2026-10-06 01:47 · r/LocalLLaMA
    Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra
  3. 2026-10-05 23:48 · r/LocalLLM
    RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to
  4. 2026-10-05 14:40 · r/LocalLLM
    Uncensored models: what they are and how they work, with a file-by-file hash check of one uncensored Qwen3.8-Flash-Next upload against Qwen's original
  5. 2026-10-04 20:23 · r/LocalLLM
    Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram

More stories

  1. Local AI ecosystem overview — r/LocalLLaMA
  2. SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE — r/LocalLLaMA
  3. Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset — r/LocalLLaMA
  4. I made a free all-in-one LoRA trainer for consumer GPUs (Windows + Linux): Qwen-Image 2.1, FLUX.2 Klein 9B, Krea 2, Z-Image, Ideogram 4, Anima, SDXL/Pony/Illustrious, LTX 2.3 and MiniMax-H3 (video + audio) — from 4-8 GB VRAM — r/StableDiffusion
  5. Qwen Flash Next on Single B200 or B300, any pointers ? — r/LocalLLM
  6. I mapped every major Qwen release from 2023 to 2026: 44 models, from Qwen-7B to the 2.4T open weights (with sources) — r/machinelearningnews
  7. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  8. LLM Inference Dashboard — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →