The operating layer that makes a fleet of AI coding agents safe to point at production systems — one canonical source per capability, cross-provider adversarial review, isolated execution, and a record of the work it stopped.
Coding agents are fast and confident. Pointed at a production system, that combination is the hazard, not the benefit. Three failure modes show up immediately and none of them are solved by a better prompt:
This is the infrastructure I built to close all three, running live across two agent runtimes and pointed at a production system carrying real revenue.
The useful measure of a review system isn't how much it approves. These are cases where the infrastructure overruled the plan — including my own.
Every team adopting coding agents at scale hits these problems in roughly this order.