Optimal 1.25 bit quantization of Qwen3.8-Flash-Next
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if: a) it would be worth it to attempt this method (since they had papers detailing it) for Qwen3.8-Flash-Next b) it would be worth my time…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-31 01:29 · r/LocalLLaMA
Optimal 1.25 bit quantization of Qwen3.8-Flash-Next