Up = more scope. Real EHR tasks (MedHELM), real clinician chats (HealthBench Professional), composite clinical
suites (MAST), the occasional RCT. The response to saturation is realism.
Not right = same cadence. Curate, freeze, grade, publish. HealthBench Pro took 15,079 real conversations
and froze them into a static test set.
The benchmark community says it itself. ARISE's 2026 report: traditional QA scores are saturated; the field
needs "multi-turn unstructured real-world data" and "prospective and post-deployment real-world scenarios."
Benchmark design is converging on watching real usage. What still differs is cadence and data access — the chasm.