AINewsnow

M5U base 96GB inference numbers for Q3.8FN after 112M tokens

TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted So the good news is that I got the base model on launch day with only 64 core GPU. All benchmarks are for current maxed out model, which looks very sw…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-23 18:40 · r/LocalLLaMA
    M5U base 96GB inference numbers for Q3.8FN after 112M tokens

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  3. Create your own voices with Gemini 3.8 text-to-speech — Google DeepMind YouTube
  4. Google releases Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its "most expressive audio generation models yet", with support for more than 100 languages (Google) — Techmeme
  5. Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge AI
  6. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  7. No Shirt, No Shoes, No Service: Amazon Blocks Meta’s Muse AI From Shopping — CNET AI
  8. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech

Get the daily brief of stories like this at 6:30 every morning →