AINewsnow

Qwen3.8 Flash Next llama.cpp config tuning

This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.

Hola all. Do you guys mind sharing your LLama.cpp config and system setup details for Qwen3.8 Flash Next? Model's quite big and tryining many combinations of llama.cpp options takes lots of time, so looking for other people setup details. I've attached my current config at the bottom, so if anyone…

Read the full story at r/LocalLLaMA ↗

Timeline · 4 reports

  1. 2026-09-14 04:30 · r/LocalLLaMA
    Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra
  2. 2026-09-12 21:08 · r/LocalLLaMA
    Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
  3. 2026-09-12 15:27 · r/LocalLLM
    Qwen3.8-Flash-Next on a 48GB M5 Pro MBP: ~20 tok/s in chat, but slow prefill makes agent use unpractical
  4. 2026-09-12 08:18 · r/LocalLLaMA
    Qwen3.8 Flash Next llama.cpp config tuning

More stories

  1. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/huggingface
  2. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  3. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  4. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  5. M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context — r/LocalLLaMA
  6. I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8... — r/LocalLLaMA
  7. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  8. My Version of Jev running locally, playing doom. — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →