Skip to content
← Articles

AI Marketing Is Not Autopilot

Revised 3 min read

By Jake Bauman

ai-systems / strategy / build-notes

The pitch is seductive. Drop an AI agent into your marketing stack and it figures out the rest. Personalizes emails. Writes copy. Runs A/B tests. All while you sleep.

From my work with marketing workflows, the surrounding decisions matter as much as the model: what it may access, what a good result looks like, and who approves external actions. Anthropic's agent design guidance also recommends starting with simpler workflows and adding autonomy only when it improves measured outcomes.

The onboarding problem

When you hire a junior marketer, you do not hand them the keys to your Klaviyo account on day one and walk away. You give them context. You show them what good looks like. You set guardrails. You review their work.

An AI agent is the same. It arrives with general intelligence but zero context about your customers, your brand, your segmentation logic, or the exceptions you learned about the hard way.

In my experience, it helps to invest in the decisions around the system: the prompts, the evaluation criteria, the feedback loops, and the holdout tests that tell them whether the agent is actually moving the needle or just generating plausible-looking output.

Three things that matter more than the model

Decision frameworks. Before the agent writes a single email subject line, you need to define what a good subject line looks like for your audience. Not "engaging" or "high-converting." Those are feelings, not criteria. Something like: specific benefit, under 40 characters, no question marks, matches the tone of the last five emails that beat baseline.

Feedback loops. The agent sends a campaign. What happens next? If the answer is "we move on to the next one," you are running automation, not a system that learns. The difference is whether the outcome of that campaign changes what the agent does tomorrow. Use the outcome to decide what to keep or change in the next run. A feedback process is useful, but it is not proof of a defensible moat.

Holdout logic. You cannot evaluate an agent the way you evaluate a static campaign. A control or holdout group can help estimate the effect of the agent-assisted work, if the test is designed for the decision and respects customers. Compare cohorts carefully and record other changes that could explain the result.

Start small, stay close

Start with one workflow, such as drafting welcome-series subject lines. Define the criteria, compare drafts with the current process, and require review before sending. Expand only after the test shows acceptable quality and a useful customer outcome.

What I am watching

Three practices I keep watching:

  • Decision ledgers. The ability to trace every agent decision back to the criteria, prompt version, and feedback data that produced it. Not just for debugging. For trust. Customers will eventually ask whether a human or an agent wrote the email they received, and you will want a good answer.

  • Evaluation-first systems. Most AI tools are generation-first: they produce, then you check. The better pattern is evaluation-first: define what success looks like, generate against those criteria, then verify. This is harder to build but produces better output.

  • The operator premium. As AI marketing tools become commodity, the scarce resource shifts from "can generate content" to "knows how to direct, evaluate, and improve the generation." People who understand both marketing and the limits of the system are better placed to make those decisions.

If you are building in this space, I would love to hear what you are learning. My inbox is open.

For a bounded starting point, see the free Growth Agent Starter Kit.

Get the free Starter Kit.

Enter your email to see the download link on this page. Build Notes is optional and has a separate checkbox.

Your files appear here after signup. If email delivery is available, you may also receive your requested file by email. The newsletter is optional. Privacy.

Related reading