AINewsnow

クロード・ハイク 5.5 ベンチマーク

This story is from 2026-10-08. It is preserved in the archive; the latest stories are on the live feed.

Claude Haiku 5.5は、OSWorld 2.1で72.4%、GDPval-AA v2.1で1620、FrontierCode 1.1で46.4%、Terminal-Bench 4.0で39.2%のスコアを記録しました。Haiku 4.5は、これら3つのうち、15.7%、735、0.0%のスコアでした。これらはすべて最大努力の数値であり、デフォルトの medium 努力では、GDPval-AAは1277に低下します。現時点では、独立した研究機関からHaiku 5.5の結果はまだ発表されていません。 今すぐApidogを試す 以下では、Claude Haiku 5.5のベンチマーク、…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-08 04:15 · DEV Community — AI
    クロード・ハイク 5.5 ベンチマーク

More stories

  1. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  2. Introducing Mistral Large 4 — Mistral AI News
  3. OpenAI will watermark ChatGPT outputs by default—but only in the EU — Ars Technica AI
  4. Claude Pro vs ChatGPT Plus vs Copilot Premium: which one would you choose for this use case? — r/ChatGPTPro
  5. Claude Haiku 5.5 now available on AI Gateway — Vercel Blog
  6. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
  7. New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent — AWS Machine Learning Blog
  8. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →