HauHauCS 27B Uncensored port to Ninfer-4090
Working this tonight. 262K context instead of 131K, and about 130 t/s decode instead of the 91.6 t/s measured on the current IQ4_XS entry, plus NInfer's faster INT8 prefill. NInfer quantizes to about 17 GB, roughly the same size as the IQ4_XS, and going through Q8_0 first adds almost no error. CPU…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-07 21:08 · r/LocalLLM
HauHauCS 27B Uncensored port to Ninfer-4090