Build AI products

AI evaluations for product managers: define good before you ship

Create a small evaluation set with realistic tasks, failure cases, a scoring rubric, and release criteria that reflect user consequences.

A convincing demo shows that an AI product can work on one example. An evaluation asks how it behaves across examples that matter, including situations where it should ask a question, refuse an action, or admit that evidence is missing.

PMs can help define these evaluations because they understand the intended task and the consequences of failure. Engineers and domain experts help make the tests reliable and technically meaningful.

Define success at the task level

For a support-drafting assistant, success is not “writes fluent text.” The draft must answer the customer's actual question, use the correct policy, preserve relevant details, and avoid making commitments the company cannot honor.

Write unacceptable outcomes separately. A draft that exposes another customer's information must fail even if its tone and grammar are excellent. A weighted average should not conceal a critical failure.

Build a small diagnostic set

Collect approved examples representing ordinary tasks, ambiguous requests, missing information, conflicting documents, and disallowed actions. Include unusual but consequential cases, not just the most common inputs.

A starter set of 20–30 cases can help a team discover obvious weaknesses. It is a practice starting point, not a statistically sufficient certification or a release guarantee. Expand it as you learn about real usage and risk.

For each case, store an ID, input, available evidence, expected behavior, unacceptable behavior, and a scoring rubric. Record the source of the example and whether it is real, modified, or fictional.

Write a rubric reviewers can apply

For a policy-answering assistant, evaluate evidence correctness, answer relevance, required caveats, and whether the next action is appropriate. Some checks can be automatic, such as whether a required field exists. Others need human judgment.

Have two reviewers independently score a subset, then discuss disagreement. If people cannot apply the rubric consistently, clarify the criteria before using its results to compare models.

An AI judge can help with scale, but validate it against human-reviewed cases. A second model is not automatically an independent source of truth.

Test the whole system

A correct response can still come from a bad workflow. Inspect which documents were retrieved, whether access rules were honored, what tools were called, and whether any action changed external state.

Anthropic's evaluation guidance explains why agent evaluation involves behavior across multiple steps. The practical implication for PMs is to inspect both the answer and the route to that answer.

Record the model, prompt, retrieval setup, tools, configuration, and test date. Repeat relevant cases because outputs can vary. Compare candidate changes against the same baseline and keep a holdout set that was not used to tune the system.

Connect evaluation to a release decision

Define gates appropriate to the task: minimum acceptable quality on critical segments, no known critical failures in the tested set, acceptable latency and cost, and a recovery path when things go wrong. Passing a finite test set does not prove that failures cannot occur.

After launch, review real failures and add representative cases to the suite. Track task success, corrections, and escalation alongside offline scores. A product can improve on a test set and still disappoint users if the task or population changes.

Try it: Use the AI Evaluation Starter Kit. Begin with one task and a rubric your team can explain.