Your AI bill doesn't tell you what you think it does

Six webinar takeaways from Faros CEO Vitaly Gordon on measuring AI spend by cost per outcome, setting smarter quotas, and choosing models using your own code.

Illustration showcasing a token bill being examined closely.

Your AI bill doesn't tell you what you think it does

Six webinar takeaways from Faros CEO Vitaly Gordon on measuring AI spend by cost per outcome, setting smarter quotas, and choosing models using your own code.

Illustration showcasing a token bill being examined closely.
Chapters

AI agents have changed the economics of software development

A large AI bill can mean a team is stuck in expensive loops, or it can mean that team is shipping faster than it ever has.

The invoice often looks the same either way.

The latest AI Engineering Report found a sharp rise in pull requests merged with no review at all as AI-generated code has scaled. That's a symptom of the same problem: volume went up, and the organization lost track of whether the volume was worth having.

Our CEO Vitaly Gordon, who co-founded Salesforce Einstein and ran its engineering org before starting Faros, recently talked through how engineering leaders can tell those two situations apart.

Here's what stuck with me during our recent webinar discussion.

One engineer, $20,000 in monthly token spend

One of our engineers spent $20,000 on tokens in a month. Nobody could say whether that person should be promoted or get a stern talking to, because the number carried no information about whether it was waste or the best money the company spent that month.

Vitaly's point is that spend only means something next to what it produced: a merged pull request, a resolved ticket, a shipped feature. We call that pairing cost per outcome. A dashboard shows where the money went and a router lowers the cost of a single task, but neither says whether the task was worth doing.

Outcomes per dollar, not tokens per dollar

Elastic engineering capacity doesn't remove the constraint on a business; it moves it.

The team burning the most tokens and the team producing the most value aren't always the same team, and a company-wide total can't show which is which.

One customer loaded their existing usage data into our platform, added nothing new, and mapped it onto their org chart. A non-engineering group was spending more on AI than the engineering organization it supported. What to do about that was the leader's call. Faros makes it visible.

Quotas cap your best people first

This was the sharpest point from our conversation. The employees who reach usage limits fastest tend to be the ones getting the most out of AI, so a flat quota lands hardest on exactly the people you would want to keep pushing. Approving every override request only postpones the question of who has earned more spend, and what evidence says so.

We build quotas around throttling. A team nearing its limit gets rerouted to a cheaper model and keeps working, so the decision becomes which route it runs on near the ceiling.

Dashboards and routers hit a limit of their own. A router can lower the cost of one task and a dashboard can show where the money went, but neither can say whether the work was worth doing.

For scale, median companies spend about $13 per engineer per month on AI. The top 1% spend around $7,000, and by most available measures they are also growing faster.

‍

A new tool doesn't fix an old process

Vitaly's comparison: hand a furniture maker a chainsaw and the shop gets no faster if nothing else about it changes. The assembly line, not the steam engine, drove the Industrial Revolution.

We've watched teams reach 70 to 80 percent AI-generated code, move every engineer onto full-time review of it, and finish slower than before AI. Cloud adoption went similarly. A credit card and an account for every developer didn't make a company cloud-native. Access was never the hard part; restructuring the work was. A general leaderboard says little about your code, so we test each model-and-harness combination against a customer's own repositories.

‍

‍

Skill varies too much for one policy to fit everyone

AI ability differs sharply between engineers, so any single company-wide quota is too generous for some and too tight for the people doing the most with it.

In Faros you can click from a person's aggregate spend into the pull requests behind it. A $2,000-per-outcome average might come from a hard repository or from work that never needed AI, and the pull requests are the only way to tell which.

Today's evidence bar won't hold up next quarter

Three major model releases came out the day this conversation was recorded. Re-evaluating each one by hand isn't realistic, so we built Time Machine. It re-runs work that already shipped, from your own history, on a different model or harness and scores the result against the original.

Across 211 real engineering tasks, the winning route held frontier quality, ran faster, and cost roughly half as much as the next-closest option.

Where to go from here

The full conversation covers more ground than six takeaways can, including how Vitaly thinks about budget conversations with finance and what changes once AI spend is treated as a capital allocation problem rather than a cost to minimize. Watch the full webinar on YouTube.

If you want to see what cost per outcome looks like against your own repositories rather than a hypothetical, book a demo and we'll walk through it on your data.

Chase Norton

Chase Norton

Chase is the Head of AI at Faros.

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration 
Whiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
AI Industry
6
MIN READ

What is a work restart? Why AI is driving them up 66.7%

A work restart is a task sent back to development after review, QA, or deployment. Learn what restarts cost in developer time and AI tokens, and how to reduce them.

AI Industry
10
MIN READ

Comprehension debt: When AI speed outpaces human understanding

Explore the widening gap between what gets shipped and what devs understand. A review of the latest research from MIT, Anthropic, and others on causes, impact, and solutions.

AI Industry
12
MIN READ

What is a software factory? How it works

Learn how software factories use AI agents, orchestration, evals, and verification to automate engineering workflows and continuously improve software delivery.