AI feature usage is not the same as value. A person may generate five answers because the first four were unusable. A dashboard that counts generations could celebrate a frustrating experience.
Start with the customer task, then measure whether people reach the intended outcome and what it costs to help them get there.
Define the unit before the metric
Choose what counts as a task, a user, a successful result, and a time window. For a meeting-to-action-plan tool, a task might begin when an eligible note is submitted and end when an approved plan is saved or the attempt is abandoned.
Count retries consistently. Decide whether a repeated request is part of the same task or a new task. Include failed and abandoned tasks where relevant so the denominator does not hide difficulty.
Build a small metric set
| Metric | Definition | What it can tell you |
|---|---|---|
| Activation | Eligible new users reaching a defined first-value event ÷ eligible new users in a specified window | Whether newcomers reach the promised initial value |
| Task success | Successfully completed eligible tasks ÷ attempted eligible tasks | Whether the workflow works in practice |
| Correction rate | Reviewed outputs requiring material correction ÷ reviewed outputs | How much quality work is transferred to users |
| Cohort retention | Members of a starting cohort returning to the value event in a defined later window ÷ eligible starting cohort | Whether value repeats at the natural task frequency |
| Escalation rate | Tasks routed to human support ÷ eligible tasks | Where users need help or the system needs a fallback |
| Cost per successful task | Defined operating cost for all task attempts ÷ successful tasks | Whether useful output is economically sustainable |
| Response latency | Time between the defined start and useful response; inspect median and tail values | Whether waiting harms task completion |
These measures require instrumentation and clear definitions. None is a universal target. A monthly planning tool should not be judged by daily retention alone.
Use a worked example
Imagine a pilot with 1,000 attempted tasks, of which 700 meet the pre-agreed success definition. Task success is 70%. Suppose all attempts together cost $120 in model and retrieval charges and $180 in variable human review. Cost per successful task is $300 ÷ 700, or about $0.43.
Dividing by all attempts gives $0.30 per attempt, which answers a different question. Neither number includes fixed engineering costs unless you explicitly add them. State the scope instead of calling the result total profitability.
If 160 of 800 reviewed outputs require material correction, correction rate is 20% among reviewed outputs. If review was not representative, do not generalize that result to every output.
Compare against a baseline
Measure the current process or a simpler alternative. If an AI workflow takes two minutes to generate and eight minutes to correct, compare that full ten minutes with the existing task. Include quality and user confidence, not only speed.
For causal claims, use a suitable experimental design. A randomized experiment can help when feasible, but sample size, assignment, duration, and guardrails matter. A before-and-after increase may reflect seasonality, different users, or another product change.
Connect metrics to decisions
Define the action each signal supports. Low activation may suggest unclear setup. High correction may require better evidence or a narrower task. Strong retention with rising costs may require a different usage limit or implementation.
Review results by meaningful segment. A healthy overall average can hide a poor experience for a specific language, user role, or input type.
Try it: Use the Metrics and Experiment Plan to define your denominator, baseline, guardrails, and next decision before launch.