Part IV

The instrument

The evaluation that fits the shape of the problem — and what each audience does about it on Monday.
12 · The thesis

Most model evaluation should be product analytics

The benchmark is the unit test; the deployment funnel is the exam. Six layers, two epochs:
Observed live (usage analytics) 1 · Delivery — did the AI produce and ship anything?
2 · Exposure — delivered ≠ viewed; who never opens it?
3 · Engagement — utilization per unit of real work
4 · Action — accept/reject from the system of record, deduplicated, correctly attributed
Reconciled after the fact (post-usage) 5 · Trust — rejection reasons as expert-labeled error analysis
6 · Outcome — evidence in someone else's system, attributed back to what the AI said

Foundation — data hygiene: exclude internal and impersonated sessions, one clock, identity resolution, one event = one decision. Uncorrected, these bias every number with a flattering sign.