Running AI marketing agents: the checks that mattered
Revised 3 min read
By Jake Bauman
ai-agents / production-readiness / evaluation / field-notes / build-in-public

These are my operating notes. I have not published an independent audit, run log, or cost ledger for the specific agents discussed here. Examples describe how I think about the work; they are not a cross-company study.
The hard part of an agent workflow is knowing whether it did good work, not simply whether a scheduled task completed. A draft can have the right format and still contradict a prior decision, cite a nonexistent source, or miss a customer constraint.
LangChain's 2026 State of Agent Engineering survey, based on more than 1,300 respondents, found that 32% cited quality as a top production barrier. That figure describes survey respondents. It does not give a failure rate for all AI agents.
The metric that misleads
Tasks completed, tokens consumed, and posts drafted tell me whether a workflow ran. They do not tell me whether the work is accurate or useful. For content, I check the factual claims, the fit with prior positions, whether a human can edit it, and whether the published page answered a real question.
I once found a draft that passed format checks but argued against a position I had taken earlier. A check for word count would never have caught it. That led me to review the substance of customer-facing work before release.
Four layers of review

Deterministic checks. Code can verify that a required field exists, a link resolves, an image file is present, or a draft stays within a defined format. These checks are cheap and repeatable. They do not prove factual accuracy.
Trace review. A useful trace shows what input the agent received, which tools it called, what those tools returned, and the final output. Sample successful and failed runs. Ask whether the agent had the right evidence and whether a handoff lost important context.
Human approval. Put a person at the boundary where an agent's work reaches a customer or changes an external system. The review should have explicit criteria and a named owner. A vague request to "take a look" often becomes a rubber stamp.
Outcome monitoring. After release, compare what happened with the goal set beforehand. For a page, indexing is one check, but qualified visits and useful actions matter more. A lack of traffic on its own does not prove that the content is poor; distribution, timing, and search intent also matter.
Problems I watch for
A weak brief can produce a polished but irrelevant answer. I now name the audience, the decision, the source material, and the expected output in the brief. A good agent cannot infer a business requirement that never reached it.
Model and tool behavior can change. I keep representative examples and rerun them after a meaningful change. If the output shifts, I inspect the reason before widening the agent's authority.
Multi-agent handoffs add another failure point. The next agent may receive a summary that omits a caveat. I prefer handoffs with required fields, source links, and a way to stop when evidence is missing. Anthropic's agent design guidance recommends adding complexity only when it improves measured outcomes. That is the right test for adding another agent.
Start with one workflow
Choose a task with a clear output and a person who can judge it. Write five good and five bad examples, including a plausible but wrong one. Add simple checks, inspect several traces, and keep approval before any external action. Measure review time and failure cases alongside task completion.
An earlier version of this article cited an unverified 847-deployment analysis and precise failure and cost figures. I removed them because I could not substantiate the underlying data. The framework above is an operating approach to test, not a promise that a particular percentage of failures will be prevented.
If you want a repeatable output review step, see the Quality Gate.