AINewsnow

4x5090 and only like 50tok/s running 320B MoE

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

I keep seeing people saying they get super fast speeds on various hardware but I can never get anything decent. I'm running GLM-5.3-Flash-UNCENSORED-FP8-UD-Q2_K_XL-00001-of-00003.gguf on a 4x5090 rig and can barely get 50tok/s. Am I missing something here? Even on larger systems like 4xH200's I was…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-11 22:12 · r/LocalLLM
    4x5090 and only like 50tok/s running 320B MoE

More stories

  1. Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators — Pandaily
  2. Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8) — r/LocalLLaMA
  3. China’s Z.ai raises revenue target 25% after US$5 billion cash injection — South China Morning Post Tech
  4. Inside ZCode (Made by GLM team): Silently Uploading Your Entire Git History to the Cloud — r/LocalLLaMA
  5. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  6. Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
  7. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  8. Optimizing agent system prompts with Amazon Bedrock AgentCore — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →