Inside Faros Token Engineering: AI Spend Observability

Go beyond token usage with AI spend observability that shows where spend goes, what work ships, and the cost per verified outcome across teams.

“Observe” displayed in bold black text on a white button against a red background.

Inside Faros Token Engineering: AI Spend Observability

Go beyond token usage with AI spend observability that shows where spend goes, what work ships, and the cost per verified outcome across teams.

“Observe” displayed in bold black text on a white button against a red background.
Chapters

The missing layer between token spend and outcomes

An AI leader at a Fortune 500 financial services firm recently told us, “I think it’s evil that the industry came up with the word token, because I have flashbacks of going to an arcade and buying tokens with no idea of how much I was spending.”

That person is not alone in their experience. Uber exhausted its entire 2026 AI budget by April, and its president said the link between AI spend and shipping products customers actually want is "not there yet." Meta ran an internal token consumption leaderboard, and total usage hit 60 trillion tokens in a single month.

All three were operating with usage data. None had observability. Usage data tells you how many tokens were used. Observability tells you what happened, to whom, on what work, and whether the outcomes were worth it. That distinction is the subject of this post. 

Why AI spend is hard to see

Traditional software costs based on seats and server usage are relatively easy to forecast quarter to quarter. AI coding spend behaves nothing like that, for three reasons:

  1. Consumption is behavioral. Two engineers with identical licenses can differ by an order of magnitude in what they cost. One might be running agents against a large monorepo and the other might be accepting autocomplete suggestions.
  2. Costs are variable. A single agentic task triggers a cascade of model calls, retries, context loads, tool invocations. Agents fail for reasons that have little to do with the model, and each failure is a retry loop.
  3. Visibility is fragmented. Your organization is likely running several AI coding tools, each with its own visibility dashboard and unaware of the others.

None of that shows up on an invoice. An invoice shows a total, but not the behavior behind it.

Spend isn’t the problem

Most organizations can already tell you what their AI spend was last month, maybe by vendor. But a chart of rising spend is meaningless on its own. What really matters is whether that spend is producing results. 

Faros operates with a concept called an outcome: a unit of work that moves the business forward. For an engineering organization that usually means a merged pull request, a resolved ticket, or a shipped feature. Once you have outcomes, you can calculate and observe cost per outcome, the only metric that matters.

But spend doesn't tell you how much AI your organization actually uses. Some people are on personal plans. Some teams have generous subscription ceilings. Some tools are still all-you-can-eat, while others have moved to metered pricing.

That cuts both ways. You may be consuming far more than your bill suggests, so a pricing change lands as a shock. Or you may be paying for capacity nobody touches. As providers move to consumption pricing, that gap becomes a live risk.

Observability means holding both numbers at once: what was consumed, and what it cost.

Questions leaders should be able to answer

When observability is working, you should be able to answer these in minutes:

  • Which teams and workloads are driving spend?
  • Which models are growing fastest, and is that growth deliberate?
  • Where are usage patterns changing?
  • Which costs are concentrated, anomalous, or unallocated?

What Faros does

Spend next to the org chart. Faros maps AI spend to the structures that already mean something to you: teams, projects, org hierarchy, repositories, Jira initiatives, and epics. You can move between those views, and they are populated from the systems you already run.

Faros recently ran this analysis for a customer who already had their usage data in hand — and the results surprised everyone. We hung the same numbers on their org chart. A non-engineering group was burning more than the engineering team it existed to support. Nobody had spotted that, because until the spend sits next to the org chart, there’s no way to even ask.

Faros maps AI spend to teams, projects, and organizational structure, revealing where costs concentrate and highlighting unexpected outliers.

People and machines, side by side. Person spend measures the AI consumption attributable to each engineer, but swarms of agents can also drive significant spend. Faros pulls both and shows them together, so you can compare what a project's people consume against what that project's automated workloads consume.

Faros lets you compare AI spend from people and automated workloads, showing total project consumption alongside the outcomes produced.

Outcomes you can click into. A cost-per-outcome number is only trustworthy if you can interrogate it. Faros lets you open the underlying pull requests and tickets, so when an engineer shows an average of $2,000 per outcome, you can look at what those outcomes actually were. Sometimes they are genuinely complex changes on a difficult repository. Sometimes they are small pull requests that aren’t justified.

Faros connects AI spend to verified outcomes, showing cost per outcome with access to the underlying pull requests and tickets.

Efficiency, and then strategic importance. Spend efficiency asks how much you get per dollar. That is necessary and not sufficient, because a team can be highly efficient at producing things nobody needed. Strategic importance asks whether it’s the right work, a much harder question. A large volume of cheap outcomes on an internal side project is not the same as a smaller volume on the release shipping next week, even when the efficiency numbers look identical.

Faros maps AI investments by efficiency, strategic importance, and total spend to show you where to cut back, fix, hold steady, or scale.

Why this is hard to replicate. None of the above comes from AI vendor data alone, because AI vendor data cannot see your repositories, your tickets, or your org chart. It comes from joining spend to the engineering systems around it. That's what Faros offers.

Observability is how you think about spend

You cannot optimize what you cannot attribute. You cannot govern what you cannot measure. Every downstream decision, such as which routes to validate, which teams to give more budget, which policies to enforce at the gateway, depends on having an explainable spend story first.

Once you can see where spend goes and what it produces, the next question becomes obvious: is this the best route for this work? That's optimization, covered in the next post. After that comes governance, turning validated decisions into policy that actually holds. 

Token engineering isn't a single skill or a one-time fix. It's the ongoing discipline of seeing, routing, and governing every dollar an organization spends on AI, so spend and outcomes move together instead of apart.

See the outcomes your AI spend is actually producing. Talk to an expert or get started with a demo.

{{cta}}

Natalie Casey

Natalie Casey

Natalie is a software engineer, and most recently—a forward-deployed engineer at Faros.

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration 
Whiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
Product
4
MIN READ

From token maxxing to outcome maxxing

Stop guessing if AI token spend pays off. Faros Token Engineering traces coding agent usage to shipped features, incidents resolved, and ROI.

Product
7
MIN READ

Inside Faros Token Engineering: AI Route Optimization

Optimize AI coding by routing each task to the right model and harness, using benchmarks from your own merged code to reduce cost per verified outcome.

Product
7
MIN READ

Inside Faros Token Engineering: AI Coding Governance

Govern AI coding across every agent and harness with policies for approved models, quotas, routing, and alerts that control spend without blunt usage caps.