Don't Sleep on EXL3 Quants
This story is from 2026-08-30. It is preserved in the archive; the latest stories are on the live feed.
I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8_O KV cache. It's a joy to use a dense 30B model at this size and still get ~30 tok/s on a VRAM-constrained laptop. It's supposed to be only slightly worse than the official 17GB K-quant at a…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-30 13:15 · r/LocalLLM
Don't Sleep on EXL3 Quants