| USMLE / MedQA era | — | Medical knowledge recall, multiple choicethe 2023–24 generation; now table stakes | Exam items | Static, saturated |
HealthBench May 2025 | OpenAI | Multi-turn health conversations against physician-written rubrics5,000 conversations · 262 physicians · 48,562 criteria | Simulated | Static + "Hard" subset |
MedHELM 2025 | Stanford + partners | 121 clinician-validated tasks across 5 categories, 35 benchmarksdecision support · documentation · communication · research · admin | Real EHR data (incl. private/gated) | Quarterly leaderboard |
ARISE · MAST 2026 | Stanford Medicine network | Composite "living benchmark" of clinical benchmarksNOHARM, MedAgentBench, CPC-Bench, PhysicianBench … | Mixed clinical | Rolling refresh |
HealthBench Professional Apr 2026 | OpenAI | The tasks clinicians actually bring to a chatbot at workcurated from 15,079 real clinician conversations | Real usage, frozen | Static |
| RCTs rare | Academia | Whole-workflow effect on real decisions and outcomese.g. Abaluck et al. 2026: assistance changed deliberation, not test appropriateness | Real deployments | Years per verdict |