I expanded FreeToken's GGUF support to Qwen MoE/dense, 1-4 bit K/I quants and sharded GGUFs. Tested 35B Ornith at 47-52 tok/s on an 8GB RTX 4060 laptop
This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.
I've been messing with FreeToken since the release because the idea behind it immediately caught my attention. Getting 35B-class MoE models running interactively on an 8GB laptop GPU is already pretty wild. The problem for me was that the initial GGUF path was much narrower than the GGUF ecosystem…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-08-24 19:35 · r/huggingface
I expanded FreeToken's GGUF support to Qwen MoE/dense, 1-4 bit K/I quants and sharded GGUFs. Tested 35B Ornith at 47-52 tok/s on an 8GB RTX 4060 laptop - 2026-08-24 19:32 · r/LocalLLM
I expanded FreeToken's GGUF support to Qwen MoE/dense, 1-4 bit K/I quants and sharded GGUFs. Tested 35B Ornith at 47-52 tok/s on an 8GB RTX 4060 laptop