Perplexity says Q4_K_M repairs 89% of the damage. Measured on the full output distribution, it repairs 67%.
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
Perplexity only ever looks at one number per position: the probability the model gave to the token that actually appeared. How the *rest* of the probability mass is arranged across the other 151,935 tokens is invisible to it. Two models can have identical perplexity and behave differently the momen…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-15 08:08 · r/LocalLLM
Perplexity says Q4_K_M repairs 89% of the damage. Measured on the full output distribution, it repairs 67%.
More stories
- AI agents/automation suggestions for a solo biz — r/AI_Agents
- Getting more accurate results - personalizations — r/ArtificialInteligence
- Opti 27B: Qwen3.8-27B in 11.8 GB at 3.47 bpw, within 0.5% of FP16 perplexity and matching Q4_K_M at 30% fewer bytes. Patched llama.cpp runtime, source public, reproduce with one command — r/LocalLLM
- claude, chatgpt, kimi or perplexity subscription (end of 2026) — r/ArtificialInteligence
- Trump announces a new 'AI Force,' but says he will not 'stifle' AI — Business Insider AI
- Google's Gemini AI hacks three other companies during security test — Sky News Technology
- Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
- Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
Get the daily brief of stories like this at 6:30 every morning →