Trust in an AI product depends partly on whether people can understand what it is doing and recover when it is wrong. A friendly tone cannot compensate for an invisible action or a confident answer with no supporting evidence.
Design around the user's task, the consequence of a mistake, and the amount of review a person can realistically perform.
Decide whether the product suggests or acts
A tool that drafts a project update is different from one that sends it to customers. Make the boundary visible. Use explicit labels such as “Draft,” “Review,” and “Send approved version” where they describe the actual workflow.
Before an action, show what will happen, where it will happen, and who will be affected. Give the person an opportunity to edit or cancel. Do not make broad permission seem like approval of one narrow step.
Show the evidence people need
For an answer based on internal documents, provide accessible source references and relevant context. Do not display decorative citations that fail to support the claim.
Preserve permissions: a citation must not expose a document the user is not allowed to see. When no adequate source exists, ask for more information or explain the limitation instead of improvising a policy.
Avoid showing a precise confidence percentage unless the score is actually calibrated for the intended use. A number that looks scientific may encourage unjustified trust.
Make correction part of the normal flow
People should be able to edit the output, reject it, or return to an earlier state where feasible. Preserve their edits when regenerating an unrelated section. Explain what information will be kept and how corrections affect later work.
For a hypothetical meeting-summary tool, useful controls include correcting an owner, removing an unsupported action item, and comparing the edited version with the original note. These controls help the user inspect the result without reconstructing the entire task.
Design the difficult states first
List the moments that could damage trust: missing evidence, conflicting sources, long waits, incomplete output, denied permissions, and failed actions. Decide what the user should see and what they can do next.
A timeout message should explain whether anything was saved or sent. A partial result should identify what is missing. An unavailable source should not be silently replaced by a guess.
Try this design critique prompt:
Review this AI workflow. For each step, identify what the user believes is happening, what the system actually does, the consequence of an error, and the available recovery. Flag hidden actions, unsupported confidence, inaccessible evidence, and review steps that require too much effort. Suggest the smallest change that improves control.
Test understanding, not just satisfaction
Ask participants what they think happened, what evidence they relied on, and what they would do if the output were wrong. Observe whether they notice a deliberately flawed example in a safe test environment.
A participant saying “I trust it” does not establish that they can detect an error. Conversely, a cautious user may be using the product appropriately when the task has serious consequences.
Measure review burden
Track whether people complete the task, how much they correct, and whether they can recover. If review is so expensive that users approve without reading, the design needs attention. Narrow the task or improve the evidence rather than adding another generic warning.
Try it: Use the difficult-state checks in the Prototype Test Kit and connect them to your AI evaluation plan.