M5U base 96GB inference numbers for Q3.8FN after 112M tokens
TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted So the good news is that I got the base model on launch day with only 64 core GPU. All benchmarks are for current maxed out model, which looks very sw…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-23 18:40 · r/LocalLLaMA
M5U base 96GB inference numbers for Q3.8FN after 112M tokens