Eras stack; they don't replace each other. Each asks a more real question: item → conversation → task → usage →
outcome. The last bar is dashed: its public instruments barely exist.
TODAY
The exam era — can it recall medicine?multiple-choice knowledge · saturated at 84–90%, at/above physician level
MedQA '20
MedMCQA '22
Med-PaLM passes '22
GPT-4 aces USMLE '23
ends as marketing:
"100% on USMLE"
The rubric era — can it converse safely?physician-written rubrics over multi-turn chats
Med-PaLM long-form '23
HealthBench May '25
length-adjusted scoring '26
The task era — can it do clinical work?real EHR data · agentic workflows
MedHELM '25 · 121 tasks
MedAgentBench '25 · 70→92% in 6mo
ARISE · MAST '26
The usage erawhat do clinicians actually ask? benchmarks curated from real use
OpenEvidence 1M/day '26
HealthBench Pro Apr '26 · 15k real chats
usage ran ahead of evaluation:
~⅔ of US physicians before any
benchmark measured their questions
The outcome eradid it change what happened to the patient?
Abaluck et al. · ARISE calls for post-deployment eval '26
the era the field is asking for — its instruments
(funnels, decision streams, reconciliation)
already run daily, on the private side of the wall