# Metrics and Experiment Plan

## Product question

User task: [ ]
Hypothesis: If we [change], then [segment] will [behavior] because [evidence-based explanation].
Baseline and source: [ ]
Smallest useful comparison: [ ]

## Metric contract

| Field | Definition |
|---|---|
| Primary metric | [Name and reason] |
| Numerator | [Exact qualifying event or outcome] |
| Denominator | [Eligible population or task attempts] |
| Window | [Start, end, cohort rules] |
| Exclusions | [Bots, internal accounts, invalid events, etc.] |
| Retries | [Same task or new task, and why] |
| Quality guardrail | [Critical failure, correction, support, or trust measure] |
| Operating guardrail | [Latency, cost, or human review burden] |
| Segments | [Roles, languages, input types, or other meaningful groups] |

## Event sketch

- task_started: task ID, user/cohort ID where appropriate, timestamp, version
- output_presented: task ID, version, latency, source availability
- output_corrected: task ID, material correction flag, reason if available
- task_completed: task ID, defined success signal
- task_failed_or_abandoned: task ID, reason if observable

Collect only properties needed for the purpose. Avoid putting raw sensitive input into analytics events. Define what can and cannot be observed reliably.

## Experiment design

Comparison: [randomized / staged pilot / qualitative / observational]
Assignment unit and contamination risk: [ ]
Sample-size/duration rationale: [consult analytics; do not invent a universal number]
Predefined decision thresholds and guardrails: [ ]
Owner who can pause: [ ]

## Worked arithmetic — hypothetical

1,000 attempted tasks; 700 successful; $300 defined variable operating cost across all attempts.

Task success = 700/1,000 = 70%.
Cost per successful task = $300/700 ≈ $0.43.
Cost per attempt = $300/1,000 = $0.30.

State which expenses the cost includes. A before-and-after difference alone does not establish causation. Record what you learned and the strength of the evidence.
