AI agents have changed the economics of software development
A large AI bill can mean a team is stuck in expensive loops, or it can mean that team is shipping faster than it ever has.
The invoice often looks the same either way.
The latest AI Engineering Report found a sharp rise in pull requests merged with no review at all as AI-generated code has scaled. That's a symptom of the same problem: volume went up, and the organization lost track of whether the volume was worth having.
Our CEO Vitaly Gordon, who co-founded Salesforce Einstein and ran its engineering org before starting Faros, recently talked through how engineering leaders can tell those two situations apart.
Here's what stuck with me during our recent webinar discussion.
One engineer, $20,000 in monthly token spend
One of our engineers spent $20,000 on tokens in a month. Nobody could say whether that person should be promoted or get a stern talking to, because the number carried no information about whether it was waste or the best money the company spent that month.
Vitaly's point is that spend only means something next to what it produced: a merged pull request, a resolved ticket, a shipped feature. We call that pairing cost per outcome. A dashboard shows where the money went and a router lowers the cost of a single task, but neither says whether the task was worth doing.
Outcomes per dollar, not tokens per dollar
Elastic engineering capacity doesn't remove the constraint on a business; it moves it.
The team burning the most tokens and the team producing the most value aren't always the same team, and a company-wide total can't show which is which.
One customer loaded their existing usage data into our platform, added nothing new, and mapped it onto their org chart. A non-engineering group was spending more on AI than the engineering organization it supported. What to do about that was the leader's call. Faros makes it visible.
Quotas cap your best people first
This was the sharpest point from our conversation. The employees who reach usage limits fastest tend to be the ones getting the most out of AI, so a flat quota lands hardest on exactly the people you would want to keep pushing. Approving every override request only postpones the question of who has earned more spend, and what evidence says so.
We build quotas around throttling. A team nearing its limit gets rerouted to a cheaper model and keeps working, so the decision becomes which route it runs on near the ceiling.
Dashboards and routers hit a limit of their own. A router can lower the cost of one task and a dashboard can show where the money went, but neither can say whether the work was worth doing.
For scale, median companies spend about $13 per engineer per month on AI. The top 1% spend around $7,000, and by most available measures they are also growing faster.
A new tool doesn't fix an old process
Vitaly's comparison: hand a furniture maker a chainsaw and the shop gets no faster if nothing else about it changes. The assembly line, not the steam engine, drove the Industrial Revolution.
We've watched teams reach 70 to 80 percent AI-generated code, move every engineer onto full-time review of it, and finish slower than before AI. Cloud adoption went similarly. A credit card and an account for every developer didn't make a company cloud-native. Access was never the hard part; restructuring the work was. A general leaderboard says little about your code, so we test each model-and-harness combination against a customer's own repositories.
Skill varies too much for one policy to fit everyone
AI ability differs sharply between engineers, so any single company-wide quota is too generous for some and too tight for the people doing the most with it.
In Faros you can click from a person's aggregate spend into the pull requests behind it. A $2,000-per-outcome average might come from a hard repository or from work that never needed AI, and the pull requests are the only way to tell which.
Today's evidence bar won't hold up next quarter
Three major model releases came out the day this conversation was recorded. Re-evaluating each one by hand isn't realistic, so we built Time Machine. It re-runs work that already shipped, from your own history, on a different model or harness and scores the result against the original.
Across 211 real engineering tasks, the winning route held frontier quality, ran faster, and cost roughly half as much as the next-closest option.
Where to go from here
The full conversation covers more ground than six takeaways can, including how Vitaly thinks about budget conversations with finance and what changes once AI spend is treated as a capital allocation problem rather than a cost to minimize. Watch the full webinar on YouTube.
If you want to see what cost per outcome looks like against your own repositories rather than a hypothetical, book a demo and we'll walk through it on your data.



.webp)

