First, make the standard machine-readable
Cloudflare keeps its engineering standards in a central repository called Cloudflare Codex. Structured RFC documents classify each requirement as SHOULD or MUST and give it an owner and a life-cycle state. A new standard can begin as guidance, move through observation, and later become a control that blocks changes.[1]
The leverage comes from the preparation before the model runs. Static analysis and linters handle deterministic rules, while AI is assigned rules that require context. Since the start of 2026, the AI code reviewer has flagged nearly 230,000 deviations and withheld approval nearly 16,000 times. Those figures show scale, but they do not reveal the false-block rate, review time, or any change in production defects.[1]
The bottleneck moves from code to rules
The same rules are applied to technical designs before implementation, to code during development, and to incident reports afterwards. A lesson from an incident can therefore become a new standard with an owner and an enforcement level, reaching the next design earlier. Review spreads across the development cycle instead of waiting for one final round of comments.[1]
I think the work automation moves is upstream, from reading code to writing the standard. A maintainer still has to decide when a SHOULD becomes a MUST, who approves an exception, and when an old rule should retire. AI can scale the application of those decisions; it cannot establish their quality on its own. There is another plausible reading: a high volume of findings may reflect a noisy reviewer rather than well-specified standards. Acceptance and reversal rates would separate the two explanations.[1]
The measurement teams should ask for
A small scorecard would be enough for a team evaluating a similar system: how many AI comments were accepted or reversed, how human review time changed, and how many defects escaped into production. Adding each standard's owner and last-updated date separates model performance from rule maintenance. When a rule is stale, a stronger model merely applies an old decision more consistently.[1]
Cloudflare's most useful design choice is the division of labor: deterministic conditions stay with conventional tools, and contextual judgment goes to AI. Teams need to preserve that boundary, avoid making every new rule blocking immediately, and feed reversed comments back into the standard. A false-block rate, review-time comparison, or production-defect measure published by the end of November would help show whether the approach produces more control or better software.[1]