AINewsnow

2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison

I've been tinkering with GLM-5.3-Flash on two 64GB CMP 170HX cards for a while, and the setup is finally stable enough that I figured I'd share it. I also compared it against the Qwen3.8-Flash-Next setup I've been using on the same machine: AWQ INT4 + FP8 PLE on vLLM. Besides PP/TG benchmarks, I ho…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-07 22:59 · r/LocalLLaMA
    2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison

More stories

  1. Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads — MarkTechPost
  2. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  3. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  4. We’re using GLM-5.3 Flash instead of frontier models on a massive production codebase — r/LocalLLaMA
  5. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  6. Anthropic Subscriptions Offer 5x+ More Value Than OpenAI — SemiAnalysis
  7. GPT-6 and Intelligent UI for everyone — OpenAI News
  8. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →