AINewsnow

Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.

​ ​I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster. ​Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-20 21:14 · r/LocalLLaMA
    Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.

More stories

  1. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  2. kimi 💀 — r/ArtificialInteligence
  3. 2 TB of Cheep Pmem200 Dimms can Run Kimi K3 at tg128 ~ 1 t/s · pp512 5.6558 — r/LocalLLM
  4. Are we over-engineering AI agent workflows? — r/AI_Agents
  5. StepFun joins the frontier: a previously non-frontier Chinese lab (StepFun) released a Kimi K3-level model, 3 times cheaper per Artificial Analysis — r/singularity
  6. A Chinese AI company just connected its model to Wall Street's leading data providers — CNBC Technology
  7. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  8. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology

Get the daily brief of stories like this at 6:30 every morning →