11 · Blind spots

What no benchmark measures

Everything before the answer Supply, delivery, exposure. Benchmarks hand the model a fully-specified vignette — perfect exposure by construction. In deployment, whether output was generated, delivered, and opened is the first half of the funnel, and it lives only in the vendor's database.
Adoption Task benchmarks grade output quality, never whether anyone accepted it. The decisive click happens where analytics SDKs aren't, and vendor staff pass through the same UI as customers — attribution has to be earned.
The workload denominator Time saved per appointment-day — the number clinicians actually care about — is measured nowhere in the public literature. DAU-style metrics measure the clinic schedule, not the tool.
Case: MAST's "Do NOHARM" demo The strongest version of output grading: expert rubrics per case, multi-judge autograder, scores for safety, precision, completeness, restraint. And the evaluation ends at the sentence — it checks whether the recommendation was correct, never whether the action was executed. Recommended ≠ accepted ≠ ordered ≠ administered ≠ outcome; every arrow lives in product telemetry. "Safety: Excellent" is a property of the text, not the visit.