Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some
edit: had to change formatting as I’m on mobile noe and the table broke.. I let Codex do a quick eval on Qwen 3.8 Flash Next IQ4_XS (served via vLLM + r9v) and Qwen 3.8 27B FP8 (served via vLLM + Radiance) This was not the full eval. I only ran a reduced quick test because the full version would ta…
Read the full story at r/LocalLLaMA ↗
Timeline · 6 reports
- 2026-09-26 21:11 · r/huggingface
Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle. - 2026-09-26 19:54 · r/LocalLLaMA
Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? - 2026-09-26 13:02 · r/LocalLLM
We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it. - 2026-09-26 08:52 · r/LocalLLaMA
Best current Qwen Flash Next Q4-ish? + worth using? - 2026-09-26 03:42 · r/LocalLLM
PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill - 2026-09-25 06:32 · r/LocalLLaMA
Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some