Confirmed bolting Q8 NGram into IQ4 Qwen no speed degradation
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
This came from another thread or comment. I forgot exactly where, but the basic idea was to replace the 51B N-gram layer in Qwen 3.8 Next with a much higher precision version. Someone running a 5090 replaced the N-gram portion of their Qwen 3.8 UD Q4 model with BF16. Since I'm already running IQ4_X…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-02 18:32 · r/LocalLLaMA
Confirmed bolting Q8 NGram into IQ4 Qwen no speed degradation