The number that worries me
When the research team sent me the final draft of our Speed Trap AI Engineering Report, I did what I expect most readers will do. I found the splashiest number, a 76.3% rise in pull requests merged with no review at all, and I felt like I had the full story.
But it took a second look to locate the number that worried me more.
It was 66.7%, the increase in work restarts per developer: tasks that had already moved on to review or QA and were sent back to the start.
Paying for the same task twice
In April, this same metric was only up 13.8%, so the latest jump is nearly 5X.
Before AI, a task sent back might just mean a developer got pulled onto something else. It was as likely to be that as bad work. Our researchers now think defective work is the more common cause, with the authoring stage producing things that don't survive review or QA. The telemetry can't confirm that directly, but restarts are the biggest source of duplicated AI spend the report found.
Duplicated spend means a restarted task was paid for twice. First in the tokens that produced the wrong output, then in the tokens that produced the replacement, with the human time spent discovering the error on top of both.
That human time is the part I think engineering teams tend to undercount.
A restart is the most expensive kind of context switch, because someone has to find their place again and recover the reasoning behind every earlier decision. Daily task contexts per developer are up 44.6%, according to our data pulled from 22,000 developers across 4,000 teams. People are supervising more work while touching less of it directly, and all of those restarts interrupt an engineer in the middle of that.
Tokens are becoming the unit of work
Right now, agents open 2% of pull requests across our dataset. At the leading edge, that share is 14% and climbing every quarter. Those teams are a preview of what the rest of the industry is about to face. Every inefficiency in this report is poised to multiply when the authors become autonomous.
I believe tokens are becoming the unit in which engineering work is bought.
Every retry, every oversized change, every restarted task is metered. The organizations that win will not necessarily be the ones that spend the most or the ones that spend the least. They will be the ones that can say, for any dollar of token spend, what outcomes were produced.
Over the past year, companies have gotten serious about AI budgets, with caps and approval steps. Meanwhile, our report shows more and more code merging without anyone reviewing it.
A cap treats every token the same. The tokens burned on a wrong answer count against it exactly like the tokens that shipped a feature. That is why spend governance and change governance cannot be handled as separate problems. A restart is where one fails the other.
Why we built Token Engineering
Our report shows restarts climbing as AI adoption deepens, and it points to missing context at the start of the work as the likeliest cause. Most engineering leaders will recognize this rework on their own teams. The token spend goes quietly uncounted, because the invoice and the shipped work live in different systems.
I’ve watched this play out before. Giving every developer AWS credits did not make a company cloud-native, and buying everyone a Zoom license did not make a company remote-first. AI coding is going the same way. The licenses are easy to hand out, but building the expertise to leverage AI across a company takes time.
That expertise is what we call Token Engineering.
Intelligence now comes in a unit, the token, and the discipline is harnessing it to maximize outcomes. The measure we care about is outcomes per dollar invested, and an outcome counts when it is in production and a customer can use it. Code sitting on a laptop does not count.
Putting this into practice means doing three things well.
Observing. You follow a token from the task to the pull request, through CI, to whatever finally shipped, and a monthly total becomes something you can look at by team, by model, and by kind of work.
Optimizing. New models launch faster than anyone can test them. Our Time Machine reads an organization's own task history to find which model, harness, and thinking level does best on which kind of work, so nobody has to spend their afternoons reading frontier-lab announcements.
Governing. The usual first move is a quota, and the people who hit it first tend to be your best, the ones pushing hardest on what the company can do with AI. Limits work when they're tied to what the spend produced.
Where to go from here
Restarts are costing teams real work, and most of it is avoidable. The recommendations section of our new report is where we lay out how to stop losing it, and I'd start there. The full Speed Trap report is free to download.