−39 to −45 pts
drop from exam-style to practice-based performance across 39 medical benchmarks
JMIR systematic review, Dec 2025
95% → 34%
one system's fall from evaluation to deployment conditions
Bean et al., via CMU 2026
40–50%
accuracy on safety-critical scenarios — the worst tier of the degradation ladder
factual 85–93 · reasoning 50–60 · diagnosis 45–55 · safety 40–50
Δ = assumptions
the gap traced to implicit protocol assumptions violated at deployment: task structure, who interacts, how outputs become decisions
CMU, 2026