What Are You Actually Approving?
What Are You Actually Approving?
A short worksheet for anyone reviewing an AI agent’s work.
An agent finishes its task. You review the result and click approve.
What does that approval establish?
Use this worksheet to examine one task you already give an agent.
A worked example from my lab
I gave an agent a narrow job: find contradictions between code and documentation, correct the documentation, and open a pull request.
The agent could not merge its own changes. I thought my approval was the control.
Then I found edits inside Python files. I could not read Python well enough to confidently certify those changes.
Here is how the control changed:
| Question | What I found |
|---|---|
| What was I approving? | Changes presented as documentation corrections |
| Could I evaluate every change? | No. Some were inside Python files |
| What evidence did I need? | A check that those edits did not change program behavior |
| What enforced the boundary? | A checker compared syntax trees before and after, allowing documentation differences |
| What happened if the check failed? | The run aborted before a branch was pushed |
My review now had evidence behind it: a machine-checked result appeared above the agent’s explanation.
I also found a separate problem. The agent had inherited access to services its task did not require, including mail and cloud storage. Preventing it from merging did not constrain everything else it could do.
I narrowed its tools and removed inherited connections.