State of Healthcare AI Benchmarks

Exams are saturated. Usage has exploded. The evaluation that answers "does this help?" happens in private, at product-analytics cadence.
Compiled August 2026 · from published benchmarks, peer-reviewed gap studies, and vendor disclosures · interactive — best viewed in a browser
Start here

Up front

One page per slide. Arrow keys (← / →) move between pages.
Part I

The record

What public evaluation actually is right now: the flagship instruments, and the five eras of question they were built to answer.
Part II

The ceiling

Why those instruments stopped discriminating — and why the field's response, moving up into more realism, is not the same as moving right into a faster loop.
Part III

Behind the wall

The evaluation that answers does this help? already exists and runs daily. It sits on the other side of a property line, and nothing it learns comes back.
Part IV

The instrument

The evaluation that fits the shape of the problem — and what each audience does about it on Monday.
Appendix

Sources