Coding agents can grade their own homework. Your marketing agent can't (yet).
So you're stuck babysitting it because you can't trust it.
Here's how to fix this:
1. Pick one subjective deliverable. Email copy, ad creative, etc.
2. Pull 50 real examples you've shipped.
3. Have your best marketer score them. Not just "good or bad." What's right, what's wrong, and why?
4. Feed that analysis into your agent to improve to improve it.
5. Have your agent create 50 outputs based on a sample of inputs.
6. Have your marketer rate and review again.
7. Keep iterating on your agent until your expert is happy.
This evals loop makes the difference between your agent giving you a sh*tty first draft that you spend more time editing than you would have spent creating it from scratch, and an AI employee who you can trust.
The Verification Gap is one of 10 trends I share in my new State of AI Report. Read it here: atherial .ai/state-of-ai