Skip to content
← Articles

Choose a primary metric for an AI agent, then watch the risks

Revised 2 min read

By Jake Bauman

ai-agents / evaluation / roi / build-notes

Editorial layout with a single gold number on charcoal

A dashboard full of metrics can obscure the decision an AI agent was meant to improve. I start by naming one primary outcome: the customer or business result that would justify the workflow. I still track quality, cost, and harm. A rising conversion rate would not make an inaccurate or privacy-violating agent acceptable.

This is a decision framework from my operating experience, not a validated claim that one metric is enough for every deployment. NIST's AI Risk Management Framework calls for measures appropriate to the use case and for testing before deployment and during operation. Its guidance is broader than financial return.

Choose the primary outcome

Write the task, the current way it is done, and what better would mean to a customer. An email follow-up agent might aim to reduce the time it takes a qualified lead to receive a useful response. A retention agent might aim to increase repeat purchases in a defined cohort. The metric should be tied to a decision you will actually make.

Record the baseline, denominator, time window, and data source. "Repeat purchase rate increased" is incomplete without knowing which customers counted and what else changed during the period.

Keep guardrails beside it

At minimum, record material factual errors, customer complaints, privacy incidents, human review time, and total operating cost. Set a stop condition for a severe error. These are not competing north-star metrics; they are the conditions under which the primary outcome is acceptable.

A workflow can appear to improve one number by selecting easier customers or sending more messages. Check the underlying segments and any holdout group before attributing an improvement to the agent.

Make the comparison credible

Compare an agent-assisted cohort with the existing process where possible. If you cannot run a control, document the other changes that might explain the result. Use a leading indicator only when it has a plausible relationship to the delayed outcome, then verify whether that relationship holds.

The previous version of this article said a four-week movement in one metric proved that an agent was working. It also used an unpublished 40% improvement example. Neither statement supports a causal conclusion, so I removed them. A useful evaluation records the number, the quality checks, the comparison, and what remains uncertain.

The free Decision Packet Writer can help capture that decision and its evidence.

Get the free Starter Kit.

Enter your email to see the download link on this page. Build Notes is optional and has a separate checkbox.

Your files appear here after signup. If email delivery is available, you may also receive your requested file by email. The newsletter is optional. Privacy.

Related reading