← LinkedIn posts

The agent ran the checks, and wrote the report saying they passed

Gatekeepers answer whether an action was allowed. They do not answer whether the agent did what it says it did. A session-end hook re-runs seven of the eight gates, so the agent's summary stops being the record.

Gatekeepers answer whether an action was allowed; they do not answer whether the agent did what it says it did, because the checks ran in the same session, scored by the same model, with no second observer. Under the hood: four hooks across three custom Claude Code agents, seven of eight quality gates re-run when a session ends, the formatter run twice in different modes, and mutation testing scoped to the changed files.
The whole post in one graphic. Open the graphic full size ↗

Your AI coding agent finished its work and reported every check green. The same agent wrote the code, ran the checks, and wrote that report.

Why it matters

There is a category of tools that stop an agent before it does something dangerous. Block the code submission. Intercept the install command. Those are gatekeepers, and they solve a real problem.

But gatekeeping answers one question: was this action allowed? It does not answer the second: did the agent do what it says it did?

An agent that runs its own checks and writes its own report has produced a claim. The checks ran in the same session, scored by the same model, with no second observer. Huang and co-authors showed at ICLR 2024, in “Large Language Models Cannot Self-Correct Reasoning Yet”, that models cannot reliably correct themselves without outside feedback. What closes that gap is a second run by something that was not in the room.

A claim is not evidence.

How it works

I run 4 hooks across 3 custom Claude Code agents. Two guard Bash commands: one blocks git operations, the other intercepts package installs. A third intercepts file writes to deliver a writing manual before the first document.

The fourth fires when a session ends. It re-runs 7 of the 8 quality gates itself and blocks the close when one fails. Six of the seven block; the type checker reports and lets you through. The agent’s summary stops being the record.

An auditor that reformats the code it inspects has changed the evidence. So the formatter runs twice, differently. In session, the agent runs Black, which rewrites files. At session end, the hook runs it in check mode: it writes nothing and exits non-zero if anything would change.

The eighth gate, mutation testing, stays out of the hook. It costs one full test run per mutant. Petrović and Ivanković, reporting on mutation testing at Google’s scale, scope it to the diff at review time. So I wrote a runner that mutates only the changed files, and it runs while the code is being written rather than at the end.

The auditor covers the two agents that write code. The research agent has no session-end check, because it produces blueprints rather than a test suite.

One bad green report in a supervised session surfaces when the code fails in front of you. The same report in an overnight batch propagates into whatever runs next, which is the case this is built for.

Where this stops: the hook holds while its script is present and executable, and a documented environment variable switches it off. That beats a prose instruction, which fails open every time the model reconsiders it. It is not a hard capability boundary, and I would rather say so than imply the pipeline is sealed.