Part II

The ceiling

Why those instruments stopped discriminating — and why the field's response, moving up into more realism, is not the same as moving right into a faster loop.
04 · Saturation

Exams stopped discriminating

84–90%
frontier accuracy on USMLE-style exams — at or above average physician performance
JMIR systematic review, Dec 2025
"100%"
a perfect USMLE score, used as consumer marketing by a clinical AI vendor
capability score as product claim
70% → 92%
best-model success on MedAgentBench in ~6 months
ARISE / MAST tracking
length-adjusted
HealthBench scoring revised after verbose answers were found to game rubric coverage
the benchmark patching its own incentive bug