Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
There's a benchmark called MedXpertQA that has difficult board level medical questions and has the questions organized by organ system. I seem to be only one to have run this benchmark against the latest models including from OpenAI/Google and published the results online. 2 years ago GPT-4o scored…
Read the full story at r/ArtificialInteligence ↗
Timeline · 1 report
- 2026-09-10 21:47 · r/ArtificialInteligence
Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM