muse glimmer architecture magic
This story is from 2026-08-19. It is preserved in the archive; the latest stories are on the live feed.
How are they able to compress context memory size? I have 32gb vram and can run llama.cpp with Q8_0 + 128k context fully without offload to system ram. It's crazy, with gemma and qwen I have to deal with lower quantization, kv quantization, etc..
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-19 16:23 · r/LocalLLM
muse glimmer architecture magic