Review overload is pushing teams to restrict AI-generated code
According to Neowin, System76, the Linux hardware maker behind Pop!_OS, recently added a mandatory checklist to the pull request template for COSMIC, its Rust-based desktop environment. Contributors must now confirm their submissions contain no AI-generated code, comments, or descriptions. The team cited review volume they could no longer absorb as the cause.
We think they have the problem right. But their remedy is only the one you pick when you have no other instruments. Below is what our own data says about how teams like System76 end up there, and what we think works better.
Our own blown spending cap
We had our own version of this.
This summer, our Claude spend reached the organization's monthly cap halfway through the month, and the cap cut off access for everyone at once. The engineer running a runaway session and the engineer shipping a careful fix lost access in the same minute, because the cap had no way to tell them both apart. We moved people to Codex. A few weeks later it happened again: 256,000 credits gone in two weeks, a 150,000 top-up, and half the team locked out on the same afternoon.
The session data was not flattering. The costliest sessions were often not the productive ones. Many ended in work that was thrown away or sent back, because the agent never knew what a reviewer would have told it: the repo's conventions, which related PRs had already tried the change, which tests had to pass. More tokens could not make up for context the agent never had. Nobody chose those defaults, and nothing told us until the limit hit.
A spending cap and a contribution ban are the same instrument. Both limit volume because nobody can see value, and the door closes on the good work along with the expensive work.
Why review capacity breaks
Our latest AI Engineering Report, The Speed Trap, is built on our own telemetry: twelve months of data from 22,000 developers across 4,000 teams, all deep into AI adoption. It documents organizations speeding up code creation while the cost of verifying and operating that code piles up downstream. These are enterprise teams, not open-source queues, where a maintainer also has to decide whether to trust a stranger, but the review mechanics are the same.
{{cta}}
In our data, average PR size is up 71.8%, and PRs merged with no review at all, human or agentic, are up 76.3%. Reviewers are seeing more than they can get through. Time in QA is up 300.6%, the sharpest deterioration of any flow metric we track, which puts the bottleneck at verification.
Sit at the end of that line with no way to rank what's coming in, and closing the intake starts to look sensible. Nothing else is within reach.
What a ban gives up
Per-change quality has stabilized. In April, teams moving from low to high adoption saw incidents per PR rise 242.7%. Among the teams in this year's dataset, already at high adoption, the figure is 14.5%, and deployment frequency is climbing again.
The aggregate is worse. Monthly incidents are up 125.4% and backlogs are growing, because each change is safer but there are far more of them. Restarts, where work is sent back to the beginning, are up 66.7%. A restart is the most expensive kind of context switch, and it shows where an agent lacked the context to get it right the first time.
A ban caps volume. It also forfeits the per-change improvement and whatever the next model generation would add.
The alternative: judge work by what it produces
If a team can't tell which contributions help, it pays either way. Most end up accepting everything and absorbing the cost downstream without ever deciding to, and the ones who refuse a category give up the work instead.
The alternative is to judge contributions by what they produce instead of how they were written. That means seeing which changes were reviewed, how long they sat in QA, what they cost, and what happened in production.
We call this token engineering: measure what AI-assisted work consumes, attribute it to what shipped, and tune model choice, context, and policy against the result.
Agentic review is the first thing we've measured at scale that moves several downstream numbers in the right direction. Teams running it on a large share of their PRs get a first review sooner and see fewer change failures. It's correlational, and unreviewed merges are still climbing even among high adopters, so it won't solve this alone.
The other place to look is before the code gets written. If an agent starts with the related PRs, the failure modes people have already hit, and the checks that matter, less of its work comes back.
What to take from a Linux distro
If you lead an engineering organization and the System76 news made you wonder whether to tighten your own policy, three moves.
- Set review requirements by risk and scope, not by author. A one-line config change and a cross-cutting refactor should not have the same rule, regardless of who or what wrote them. Decide which review applies to which class of change, and check whether the rule survives a week of delivery pressure.
- Watch time in QA as a leading indicator. It is where the Speed Trap shows up first and loudest. If it is climbing while throughput looks healthy, the pipeline is unblocked but leaky.
- Invest in what agents know before they write. Measure restarts. Trace them back to the context that was missing. The data puts the highest-leverage fix at the authoring stage, before anything reaches review.
Closing the door is a reasonable move when you cannot see what is coming through it. The better move is to make it visible.
The full Q3 AI Engineering report, with ten recommendations across observing, optimizing, and governing AI at scale, is at faros.ai/research/speed-trap.





.webp)