AINewsnow

Qwen flash next on 12+16gb vram, and 32gb ram viable?

I have 4070 12gb, and v100 16gb, and 32gb 3200mhz ram, with the basic llama, with flash attention 98k ctx on q4, with layer split NOT tensor parallelism, as the v100 is sxm2 and has a adapter with pcie 3.0 x16 and I have pcie 4.0 x4, so tensor parallelism slows this down by like .7 tps, Im on ubunt…

Read the full story at r/LocalLLM ↗

Timeline · 5 reports

  1. 2026-10-01 05:27 · r/LocalLLaMA
    Qwen Flash Next MTP work restarted
  2. 2026-09-30 17:15 · r/LocalLLaMA
    How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?
  3. 2026-09-30 14:29 · r/LocalLLM
    Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
  4. 2026-09-30 09:41 · r/LocalLLM
    Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo
  5. 2026-09-29 10:21 · r/LocalLLM
    Qwen flash next on 12+16gb vram, and 32gb ram viable?

More stories

  1. What model sits between Qwen 3.8 27b and Flash next for coding? — r/LocalLLaMA
  2. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s — r/LocalLLM
  4. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
  5. What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects? — r/LocalLLaMA
  6. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/artificial
  7. How to get desired result and consistency? — r/comfyui
  8. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →