Two levels, one decision
Effort levels in Copilot code review have reached general availability, and the list of options is down to two: Lite for a straightforward change, Balanced for a change that deserves a closer look. The pair replaces the Low and Medium levels from the public preview; the names changed, but existing configuration carries over on its own, so no team has to reconfigure anything. What sets Balanced apart is where the change goes rather than how long the review runs: it reaches a model that reasons at greater depth.[1]
Where that setting lives is the part that actually changes for a developer. An organisation administrator picks the default under Copilot code review in organisation settings, and repositories inherit it unless they choose their own level. Review depth therefore stops being a decision the person opening a pull request makes in the moment and becomes configuration the repository inherits, and it returns to the repository the moment that repository sets a level of its own. The feature is open on the Copilot Pro, Pro+, Max, Business and Enterprise plans.[1]
What the label makes testable
The least-discussed detail in the release is, I think, the most useful one: which effort level ran is now labelled in timeline events and in pull request comments. When a review misses a bug, a team can tell whether it is arguing about the model falling short or about that change having gone through on Lite. It helps to be clear about what the label does: it reports which setting applied, and carries no data on whether that setting improved the outcome. It may also help more with newly opened reviews than with a backward audit, since older reviews do not carry the information.[1]
In my July 25, 2026 piece on Opus 5, I argued that what the release brought a developer was a per-request effort and cost dial. The same dial has now moved out of the model interface and into a review product, shared this time between an organisation default and repository inheritance. In that piece I said every measurement backing the dial came from the vendor; this announcement leaves that where it was, but the label opens the first place where a team could gather data of its own.[1], [2]
Which comparison is missing?
The announcement carries no measurement of which change belongs at which level: if one set of pull requests were reviewed under Lite and under Balanced, we are not told how missed bugs, review time and the load of unnecessary comments would differ. Until such a comparison is published, the administrator setting the default is left with whatever intuition the two names suggest. There is a concrete signal teams can watch: if a study appears by October 31, 2026 that compares the two levels with the change set and the scoring held constant, and reports missed bugs alongside review time, choosing a default stops being intuition and becomes a measurable preference.[1]