Clinician use of language models diverges from how the models are evaluated
arXiv:2610.11069v1 Announce Type: new Abstract: Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CL
Clinician use of language models diverges from how the models are evaluated