Ran Gemini 3.8 Flash on an expert level medical knowledge benchmark. Scores higher than Opus 5 and GPT-5.6 Sol
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
There's a benchmark called MedXpertQA that has board level medical questions and has the questions organized by organ system. I seem to be only one to have run the benchmark against Gemini's latest models and published the results online. Interestingly 3.8 Flash scores higher than every other LLM.…
Read the full story at r/GeminiAI ↗
Timeline · 2 reports
- 2026-09-10 21:47 · r/ArtificialInteligence
Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM - 2026-09-09 21:22 · r/GeminiAI
Ran Gemini 3.8 Flash on an expert level medical knowledge benchmark. Scores higher than Opus 5 and GPT-5.6 Sol