AINewsnow

M5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds

I've been playing around with various models on the M5 Ultra 256GB 80-core Mac Studio. These are the results over many rounds of agentic inferencing. I'm happy with the performance. Glad to have the large amount of RAM. But it does feel like the GPU is underpowered for this amount of RAM. I'm wonde…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-25 02:45 · r/LocalLLaMA
    M5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds

More stories

  1. Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning — MarkTechPost
  2. We interviewed GPT-OSS, Qwen, Gemma and GLM across 24 subjects and published all 1,452 positions — r/artificial
  3. 299 real user intents tested Jev against production base line. Here is the result. — r/AI_Agents
  4. MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap — r/LocalLLaMA
  5. Z.ai disables coding assistant feature after flaw exposed enterprise code upload risk — InfoWorld AI
  6. Zhipu says ZCode removed repository-upload paths after data controversy — TechNode
  7. Zhipu Open-Sources ZCode After Repo-Upload Fix; CAICT and NSFOCUS Audits Cite Removals — Pandaily
  8. Introducing GPT-6 Sol and Luna — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →