AINewsnow

Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

There's a benchmark called MedXpertQA that has difficult board level medical questions and has the questions organized by organ system. I seem to be only one to have run this benchmark against the latest models including from OpenAI/Google and published the results online. 2 years ago GPT-4o scored…

Read the full story at r/ArtificialInteligence ↗

Timeline · 1 report

  1. 2026-09-10 21:47 · r/ArtificialInteligence
    Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM

More stories

  1. Gemini 4 Pro vs Gemini 3.8 Flash (Pelican Riding a Bicycle SVG) — r/singularity
  2. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  3. Gemini vs ChatGPT for turning research notes into a presentation structure — r/GeminiAI
  4. Zbudowalismy mala orkiestrę botow (i nie wierzymy pierwszemu google) — DEV Community — AI
  5. What are you all using for long term project/conversational memory these days? — r/LocalLLaMA
  6. One prompt two models — r/AI_Agents
  7. Thank you Google 🎉🥳🥳 — r/OpenAI
  8. Solving image to text captchas — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →