AINewsnow

GLM 5.3, GLM 5.3 Flash or 3.8 Qwen Flash for Replacing Kimi k3 IQ2_XXS

This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.

Kimi seem fine at IQ2_xxs but is slow 4tks (passable) but it can drop to 2tks (well, not great) doesn't seem so viable. Would the new GLM(s), or Qwen Flash a good substitute, especially to have a better interactive experience while having high intelligence.

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-08-31 04:38 · r/LocalLLaMA
    GLM 5.3, GLM 5.3 Flash or 3.8 Qwen Flash for Replacing Kimi k3 IQ2_XXS

More stories

  1. Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8) — r/LocalLLaMA
  2. What's the best open weight model for Blender? That's comparable to Astra — r/LocalLLaMA
  3. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  4. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  5. Post-training image models for fandom — Character.AI Blog
  6. Testing Qwen 3.8 27B running locally on a single 5090 — r/LocalLLM
  7. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  8. US government website used Chinese model the FBI called "malicious" — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →