A pilot needs a comparator, not a slogan
Measurement begins before the intervention. Define the current workflow, observation period, population, operational definitions, data source, owner, and known limitations. Without that comparator, a pilot can report usage or model quality while leaving the business outcome unknown.
Measure the operating system in layers
| Layer | Representative measures |
|---|---|
| Business outcome | Revenue protected, service resolution, capacity, loss avoided, compliance finding, customer result, or risk-adjusted value. |
| Process flow | Cycle time, wait time, touch time, throughput, backlog, handoffs, rework, exception rate, abandonment, and first-pass yield. |
| System quality | Task success, critical errors, calibration, grounded claims, citation accuracy, abstention, consistency, and subgroup performance. |
| Human operation | Review minutes, acceptance, override, escalation, queue age, fatigue indicators, training, and ability to detect errors. |
| Control and resilience | Unauthorized action, disclosure, policy failure, drift, incident detection, rollback time, availability, and recovery. |
| Economics | Build, integration, model, infrastructure, human review, monitoring, support, error loss, change, and cost per successful outcome. |
Choose one primary measure and several guardrails
The primary measure should represent the intended business result. Guardrails prevent local improvement from hiding harm elsewhere. A triage pilot may target faster time to correct routing while constraining critical misrouting, human review time, customer complaints, security events, and cost per accepted case.
Do not convert time into money automatically
Saved minutes become economic value only through a defined mechanism: reduced external spend, avoided hiring, more throughput, faster revenue, improved service, lower error, or redeployed expert capacity. The model should show conservative, expected, and upside cases and include the cost of capturing the benefit.
Measurement traps
- Comparing a carefully selected pilot group with an unmeasured historical average.
- Reporting model accuracy while excluding workflow exceptions and human correction.
- Counting adoption as value without measuring accepted outcomes.
- Ignoring the cost of evaluation, data work, security review, monitoring, and vendor change.
- Using average performance to conceal a critical error or poor result for a material subgroup.
- Changing the prompt, model, rules, or population during the final acceptance run.
Release decision
The measurement plan should be approved before the final test. The decision meeting then compares the frozen acceptance results with the baseline and guardrails and records scale, conditional scale, revise, revert, or stop.