The AI agent kept rejecting review findings by citing the rules, so I took its veto away

日本語

Contents of this article

Summary

I forbade my AI agent from rejecting the findings of a Japanese-language review by an external AI model on the ground that the text already conformed to an existing rule. On August 12, 2026, I introduced a step that sends articles about to be published to an AI model from an outside provider for review. Over the six days that followed, the agent swapped the grounds it used to reject or defer findings five times. I sent four of those back and had the agent apply what it had rejected. The fifth happened while measuring the quality of reviews, so I corrected it on the spot. The agent applied what I sent back on August 12 and 13. The text read better after it was applied. Now the practice is this: findings that arrive are treated as correct and applied to the body first, and only the ones I cannot accept come to me to decide, with the original and the revision side by side. There is exactly one exception that may be rejected, a suggestion that adds an actor or a reason not present in the actual records.

In this article I first describe how the grounds for rejection changed five times. Next I consider why the agent skipped the judgment that reads for meaning. Then I write about a different failure that came out of the accept-everything policy, and my reason for narrowing the exception to one. Finally I explain the principle I built into the operation.

What you can take away

For anyone who has an AI agent process review findings, this article covers the following three things.

In the body I refer to the external AI models that do the reviewing by their role rather than by product name. The content of this article does not represent the views of any provider.