Free resource · Measure and evaluate

Metrics and Experiment Plan

Define outcome metrics, denominators, guardrails and decision rules for an AI feature or experiment.

Product question

User task: [ ]
Hypothesis: If we [change], then [segment] will [behavior] because [evidence-based explanation].
Baseline and source: [ ]
Smallest useful comparison: [ ]

Metric contract

FieldDefinition
Primary metric[Name and reason]
Numerator[Exact qualifying event or outcome]
Denominator[Eligible population or task attempts]
Window[Start, end, cohort rules]
Exclusions[Bots, internal accounts, invalid events, etc.]
Retries[Same task or new task, and why]
Quality guardrail[Critical failure, correction, support, or trust measure]
Operating guardrail[Latency, cost, or human review burden]
Segments[Roles, languages, input types, or other meaningful groups]

Event sketch

  • task_started: task ID, user/cohort ID where appropriate, timestamp, version
  • output_presented: task ID, version, latency, source availability
  • output_corrected: task ID, material correction flag, reason if available
  • task_completed: task ID, defined success signal
  • task_failed_or_abandoned: task ID, reason if observable

Collect only properties needed for the purpose. Avoid putting raw sensitive input into analytics events. Define what can and cannot be observed reliably.

Experiment design

Comparison: [randomized / staged pilot / qualitative / observational]
Assignment unit and contamination risk: [ ]
Sample-size/duration rationale: [consult analytics; do not invent a universal number]
Predefined decision thresholds and guardrails: [ ]
Owner who can pause: [ ]

Worked arithmetic — hypothetical

1,000 attempted tasks; 700 successful; $300 defined variable operating cost across all attempts.

Task success = 700/1,000 = 70%.
Cost per successful task = $300/700 ≈ $0.43.
Cost per attempt = $300/1,000 = $0.30.

State which expenses the cost includes. A before-and-after difference alone does not establish causation. Record what you learned and the strength of the evidence.