Choose a primary metric for an AI agent, then watch the risks
Revised 2 min read
By Jake Bauman
ai-agents / evaluation / roi / build-notes

A dashboard full of metrics can obscure the decision an AI agent was meant to improve. I start by naming one primary outcome: the customer or business result that would justify the workflow. I still track quality, cost, and harm. A rising conversion rate would not make an inaccurate or privacy-violating agent acceptable.
This is a decision framework from my operating experience, not a validated claim that one metric is enough for every deployment. NIST's AI Risk Management Framework calls for measures appropriate to the use case and for testing before deployment and during operation. Its guidance is broader than financial return.
Choose the primary outcome
Write the task, the current way it is done, and what better would mean to a customer. An email follow-up agent might aim to reduce the time it takes a qualified lead to receive a useful response. A retention agent might aim to increase repeat purchases in a defined cohort. The metric should be tied to a decision you will actually make.
Record the baseline, denominator, time window, and data source. "Repeat purchase rate increased" is incomplete without knowing which customers counted and what else changed during the period.
Keep guardrails beside it
At minimum, record material factual errors, customer complaints, privacy incidents, human review time, and total operating cost. Set a stop condition for a severe error. These are not competing north-star metrics; they are the conditions under which the primary outcome is acceptable.
A workflow can appear to improve one number by selecting easier customers or sending more messages. Check the underlying segments and any holdout group before attributing an improvement to the agent.
Make the comparison credible
Compare an agent-assisted cohort with the existing process where possible. If you cannot run a control, document the other changes that might explain the result. Use a leading indicator only when it has a plausible relationship to the delayed outcome, then verify whether that relationship holds.
The previous version of this article said a four-week movement in one metric proved that an agent was working. It also used an unpublished 40% improvement example. Neither statement supports a causal conclusion, so I removed them. A useful evaluation records the number, the quality checks, the comparison, and what remains uncertain.
The free Decision Packet Writer can help capture that decision and its evidence.