14 · Sources
Sources
Benchmarks
MedHELM (arXiv)
·
MedHELM (Nature Medicine)
·
medhelm.org leaderboard
ARISE
·
MAST framework
·
MAST leaderboard
·
State of Clinical AI Report 2026
HealthBench (OpenAI)
·
HealthBench Professional (arXiv)
The gap
Knowledge–Practice Gap, 39 benchmarks (JMIR, Dec 2025)
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions (CMU, 2026)
·
blog version
Abaluck et al. — Does LLM Assistance Improve Healthcare Delivery? (NBER w34660)
Deployment-Centered Evaluation (arXiv)
·
NEJM AI clinical-reasoning benchmark
Usage
OpenEvidence — 1M consultations/day (Mar 2026)
·
~⅔ of US physicians (secondary)
Doximity 2026 State of AI in Medicine
·
press release
Companions
The six-layer monitoring timeline
·
The interactive lens quadrant
·
Essay: Model evaluation should look like product analytics
← What should happen next
15 / 15 ·
agenda
Index →