Five quality gates passed, graded by the agent that wrote the code
The agent wrote the code, wrote its tests, ran the checks, and reported five green. Nothing in the pipeline was built to ask who checked the checker.
5 quality gates on my AI coding agents. Every one of them passed. And every one of them was graded by the agent that wrote the code.
Why it matters
The agent wrote the code, wrote its tests, ran the checks, and reported the results. All 5 green. Nothing in the pipeline was built to ask: who checked the checker?
That sounds philosophical until your agent authors a test that agrees with a bug. The test passes. The gate reports a pass. The human sees green. And nobody catches it, because the only witness was the one who wrote it.
Work presented at NeurIPS in 2024 found that language models recognize their own writing and rate it higher because they recognize it. My pipeline had that bias sitting in its scoring loop and no way to see it.
An auditor must observe without mutating what it audits.
How it works
Three custom Claude Code agents: research, build, review. The first version ran 5 gates before any commit. Four confirmed syntax, style, types, and vulnerability patterns. The fifth was pytest, scoring tests the same agent had written.
A test you have never seen fail has proven nothing. That line drove the rebuild. I added three gates the first five could not see. The one that matters here is mutation testing: it breaks the source on purpose to check whether the suite notices, so it grades the tests instead of the code.
Then the harder half. When a session ends, a Stop hook re-runs 7 of the 8 gates and refuses to close on a divergence, so the agent’s report becomes an observation rather than a verdict. The formatter runs there in check mode, writing nothing: an auditor that reformats the code has changed the evidence.
Mutation testing is the eighth, and the hook does not re-run it. Full-tree mutation costs one test-suite run per mutant. Google scopes it to the diff at review time for that reason, so I wrote a script that does the same.
One engineer on one session catches a bad green report. Agents running unattended or in parallel have no such reader, and that is the case worth building for.
Where this stops: one of my three agents can still install packages unguarded. It cannot commit, push, or tag, so the blast radius is a broken environment rather than a published mistake.