AINewsnow

Dynamic Quantiser - a way to make your own high quality dynamic quants

Dynamic Quantiser TLDR - Makes dynamic and/or custom quants of ggufs. Doesn't need data, just gguf file + llama.cpp. Minimises cosine deviation - highly correlated with minimising KLD. Fast, much better than standard quants, not as good as Unsloth on pure text, maybe as good on mixed inputs/code. F…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-22 14:59 · r/LocalLLaMA
    Dynamic Quantiser - a way to make your own high quality dynamic quants

More stories

  1. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  2. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  3. Transformers now runs llama.cpp quants — Hugging Face Blog
  4. Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken — r/LocalLLaMA
  5. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  6. I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
  7. Performance tune for gemma4-26b-a4b flash attention shape. by frobnitzem · Pull Request #28450 · ggml-org/llama.cpp · GitHub — r/LocalLLaMA
  8. Is llama.cpp meant to be slow at long context, even when you aren't using that context? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →