KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy
This story is from 2026-10-01. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.38480v1 Announce Type: new Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which exami…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CL
KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy