An AI agent told the user "I filed your bug" - but nothing was filed, and the test still passed.
@pandemicsyn wrote a hands-on guide to evals, the tests that catch this kind of thing in AI agents. You work through a small demo agent that fails on purpose, and you fix the checks until they can tell the difference between an agent that did the work and one that only said it did. His coding agent skill can walk you through it if you'd rather not read the whole post.
If your test can't tell those two apart, you don't know what your agent is doing.
Read more here: