The benchmark is the unit test; the deployment funnel is the exam. Six layers, two epochs:
Observed live (usage analytics) 1 · Delivery — did the AI produce and ship anything?
2 · Exposure — delivered ≠ viewed; who never opens it?
3 · Engagement — utilization per unit of real work
4 · Action — accept/reject from the system of record, deduplicated, correctly attributed
Reconciled after the fact (post-usage) 5 · Trust — rejection reasons as expert-labeled error
analysis
6 · Outcome — evidence in someone else's system, attributed back to what the AI said
Foundation — data hygiene: exclude internal and impersonated sessions, one clock, identity resolution, one event
= one decision. Uncorrected, these bias every number with a flattering sign.