Skip to content
← Articles

Running AI marketing agents: the checks that mattered

Revised 3 min read

By Jake Bauman

ai-agents / production-readiness / evaluation / field-notes / build-in-public

Dark editorial graphic about operating AI marketing agents

These are my operating notes. I have not published an independent audit, run log, or cost ledger for the specific agents discussed here. Examples describe how I think about the work; they are not a cross-company study.

The hard part of an agent workflow is knowing whether it did good work, not simply whether a scheduled task completed. A draft can have the right format and still contradict a prior decision, cite a nonexistent source, or miss a customer constraint.

LangChain's 2026 State of Agent Engineering survey, based on more than 1,300 respondents, found that 32% cited quality as a top production barrier. That figure describes survey respondents. It does not give a failure rate for all AI agents.

The metric that misleads

Tasks completed, tokens consumed, and posts drafted tell me whether a workflow ran. They do not tell me whether the work is accurate or useful. For content, I check the factual claims, the fit with prior positions, whether a human can edit it, and whether the published page answered a real question.

I once found a draft that passed format checks but argued against a position I had taken earlier. A check for word count would never have caught it. That led me to review the substance of customer-facing work before release.

Four layers of review

Four-layer AI agent evaluation stack: deterministic checks, trace review, human approval, and monitoring

Deterministic checks. Code can verify that a required field exists, a link resolves, an image file is present, or a draft stays within a defined format. These checks are cheap and repeatable. They do not prove factual accuracy.

Trace review. A useful trace shows what input the agent received, which tools it called, what those tools returned, and the final output. Sample successful and failed runs. Ask whether the agent had the right evidence and whether a handoff lost important context.

Human approval. Put a person at the boundary where an agent's work reaches a customer or changes an external system. The review should have explicit criteria and a named owner. A vague request to "take a look" often becomes a rubber stamp.

Outcome monitoring. After release, compare what happened with the goal set beforehand. For a page, indexing is one check, but qualified visits and useful actions matter more. A lack of traffic on its own does not prove that the content is poor; distribution, timing, and search intent also matter.

Problems I watch for

A weak brief can produce a polished but irrelevant answer. I now name the audience, the decision, the source material, and the expected output in the brief. A good agent cannot infer a business requirement that never reached it.

Model and tool behavior can change. I keep representative examples and rerun them after a meaningful change. If the output shifts, I inspect the reason before widening the agent's authority.

Multi-agent handoffs add another failure point. The next agent may receive a summary that omits a caveat. I prefer handoffs with required fields, source links, and a way to stop when evidence is missing. Anthropic's agent design guidance recommends adding complexity only when it improves measured outcomes. That is the right test for adding another agent.

Start with one workflow

Choose a task with a clear output and a person who can judge it. Write five good and five bad examples, including a plausible but wrong one. Add simple checks, inspect several traces, and keep approval before any external action. Measure review time and failure cases alongside task completion.

An earlier version of this article cited an unverified 847-deployment analysis and precise failure and cost figures. I removed them because I could not substantiate the underlying data. The framework above is an operating approach to test, not a promise that a particular percentage of failures will be prevented.

If you want a repeatable output review step, see the Quality Gate.

Get the free Starter Kit.

Enter your email to see the download link on this page. Build Notes is optional and has a separate checkbox.

Your files appear here after signup. If email delivery is available, you may also receive your requested file by email. The newsletter is optional. Privacy.

Related reading