AINewsnow

Ran Gemini 3.8 Flash on an expert level medical knowledge benchmark. Scores higher than Opus 5 and GPT-5.6 Sol

This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.

There's a benchmark called MedXpertQA that has board level medical questions and has the questions organized by organ system. I seem to be only one to have run the benchmark against Gemini's latest models and published the results online. Interestingly 3.8 Flash scores higher than every other LLM.…

Read the full story at r/GeminiAI ↗

Timeline · 2 reports

  1. 2026-09-10 21:47 · r/ArtificialInteligence
    Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM
  2. 2026-09-09 21:22 · r/GeminiAI
    Ran Gemini 3.8 Flash on an expert level medical knowledge benchmark. Scores higher than Opus 5 and GPT-5.6 Sol

More stories

  1. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  2. Dumbest solution to the alignment problem — r/singularity
  3. Gemini vs ChatGPT for turning research notes into a presentation structure — r/GeminiAI
  4. One prompt two models — r/AI_Agents
  5. Solving image to text captchas — r/AI_Agents
  6. Is it just me, or does Google Gemini sometimes give better content than ChatGPT? 👀 — r/GeminiAI
  7. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  8. ChatGPT’s image generation has improved A LOT — r/ChatGPT

Get the daily brief of stories like this at 6:30 every morning →