AINewsnow

Evaluaciones comparativas de Claude Haiku 5.5

Claude Haiku 5.5 obtiene un 72.4% en OSWorld 2.1, 1620 en GDPval-AA v2.1, 46.4% en FrontierCode 1.1 y 39.2% en Terminal-Bench 4.0. Haiku 4.5 obtuvo 15.7%, 735 y 0.0% en tres de esos benchmarks. Todos son resultados con esfuerzo máximo; con el valor predeterminado medium , GDPval-AA baja a 1277. Nin…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-08 04:19 · DEV Community — Machine Learning
    Evaluaciones comparativas de Claude Haiku 5.5

More stories

  1. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  2. Introducing Mistral Large 4 — Mistral AI News
  3. OpenAI will watermark ChatGPT outputs by default—but only in the EU — Ars Technica AI
  4. Claude Pro vs ChatGPT Plus vs Copilot Premium: which one would you choose for this use case? — r/ChatGPTPro
  5. Claude Haiku 5.5 now available on AI Gateway — Vercel Blog
  6. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
  7. New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent — AWS Machine Learning Blog
  8. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →