A gate you cannot compare

Checked 22 Sep 2026 · By Luke Czak

ArticleOpinionFree to read

Tying independent review to whichever engine happens to be free that day makes every verdict a different kind of check. Independence comes from a session with no memory of the build, not from which vendor’s name is on it.

For a long stretch the rule I ran independent review on was simple: whichever engine built a change, a different engine had to grade it before the work counted as done. If one agent wrote the code, the review had to come from somewhere else entirely, a different vendor, a different set of habits. The logic felt sound. Two engines trained differently and built differently tend to miss different things, so putting a second one in the loop should catch whatever the first one could not see about its own output. It held up fine on paper, and it held up fine in the easy weeks, when both engines happened to be sitting idle and available whenever a build finished.

It stopped holding up the week both were busy at once. The honest choices at that point were to wait for one to free up, to skip the gate on something that looked small, or to let the engine that had not personally written the change review it anyway, on the reasoning that at least a second pass had happened even if it was not the pass the rule demanded. I made that last call more than once, told myself it was a one-off each time, and eventually a change went out that a stricter review would have caught, because the version of review that actually ran that day was a weaker one than every other task in the batch had received.

The deeper problem was not that one task got a worse review than another. It was that I could no longer compare verdicts across tasks at all, because the process producing them was not the same process twice. A pass from a week where both engines were free meant something different from a pass squeezed out under a rule I had quietly bent, and nothing in the record said which was which. A verdict is only worth anything if the same word means the same thing every time it is written down.

What actually makes a review independent, I eventually worked out, is not the vendor tag on the reviewing engine. It is a session with no memory of the build, holding a brief written to find fault rather than to confirm the work is fine. A builder grading its own output in the same session is the exact failure that rule was trying to prevent, and a fresh session run on the same engine that built the change blocks that failure just as completely as a different vendor would, because the thing doing the blocking was never the vendor in the first place.

So I rewrote the rule around the property that actually mattered and dropped the one that only felt like it mattered. The gate is now the same shape every time, run the same way regardless of who built the change or who happens to be idle that afternoon, which means a verdict from one task now means the same thing as a verdict from another. That comparability is worth more to me than the diversity a second vendor occasionally added, because the diversity was never reliable in the first place — it depended on both engines being free at the same moment, which is exactly the condition that broke the old rule.

I still think there was a real loss in fixing it this way. Two different models genuinely do miss different things sometimes, and a rule that always uses the same reviewing process gives up whatever that difference used to catch. I made the trade anyway, because a rule that quietly changes shape under pressure was never actually a rule. It was a preference that held only when nothing was busy, and the entire point of a gate is that it has to hold on the days it is inconvenient.

Comments (0)

Sign in to comment.