Exams are saturated. Frontier models score 84–90% on USMLE-style tests — at or above physician level. One vendor markets a perfect score.
The gap is measured. Performance falls 39–45 points from exams to practice tasks (39 benchmarks); one system fell 95% → 34% from evaluation to deployment.
Usage ran ahead. ~⅔ of US physicians use an AI tool; one platform logs ~27M encounters a month. Adoption follows peer signal, not leaderboards.
New benchmarks use real data, same cadence. Real EHR tasks, real clinician chats — still shipped as frozen test sets.
The evaluation that matters is private. Decision streams, utilization, outcome reconciliation exist only as vendor telemetry — invisible to the literature.