Part I

The record

What public evaluation actually is right now: the flagship instruments, and the five eras of question they were built to answer.
02 · The landscape

The flagship evaluations, mid-2026

BenchmarkStewardWhat it gradesDataCadence
USMLE / MedQA eraMedical knowledge recall, multiple choicethe 2023–24 generation; now table stakesExam itemsStatic, saturated
HealthBench May 2025OpenAIMulti-turn health conversations against physician-written rubrics5,000 conversations · 262 physicians · 48,562 criteriaSimulatedStatic + "Hard" subset
MedHELM 2025Stanford + partners121 clinician-validated tasks across 5 categories, 35 benchmarksdecision support · documentation · communication · research · adminReal EHR data (incl. private/gated)Quarterly leaderboard
ARISE · MAST 2026Stanford Medicine networkComposite "living benchmark" of clinical benchmarksNOHARM, MedAgentBench, CPC-Bench, PhysicianBench …Mixed clinicalRolling refresh
HealthBench Professional Apr 2026OpenAIThe tasks clinicians actually bring to a chatbot at workcurated from 15,079 real clinician conversationsReal usage, frozenStatic
RCTs rareAcademiaWhole-workflow effect on real decisions and outcomese.g. Abaluck et al. 2026: assistance changed deliberation, not test appropriatenessReal deploymentsYears per verdict