What settings do you use for running Qwen3.8-Flash-Next in llama.cpp?
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
Hi, I'm wondering what settings you are using in order to run Qwen3.8-Flash-Next on your devices? I'm especially interested in setups with 96GB VRAM. I'm not quite sure if llama.cpp does offload the embeddings to RAM or disk with my settings. I would like to offload them to RAM in order to avoid to…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-09 12:36 · r/LocalLLaMA
What settings do you use for running Qwen3.8-Flash-Next in llama.cpp?