AINewsnow

Can we please have some error bars?

I am sure this gripe has been raised many times before, but every time a new model is released it seems like it's routinely only a few percentage points higher than previous models on benchmarks. How do we know this is even a "real" difference and not just within the window of measurement error or…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-06 16:10 · r/LocalLLaMA
    Can we please have some error bars?

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  3. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  4. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  5. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  6. OpenAI safety leader quits, warning AI company’s culture is ‘broken’ — The Guardian AI
  7. Anthropic expands its Claude Startups program, including up to $45K in discounts and credits via the Claude Startup Stack and a one-time $1,000 API credit (Ashley Capoot/CNBC) — Techmeme
  8. Manage Amazon SageMaker HyperPod Spaces directly from SageMaker Studio — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →