Inside Faros Token Engineering: AI Coding Governance

Govern AI coding across every agent and harness with policies for approved models, quotas, routing, and alerts that control spend without blunt usage caps.

“Govern” displayed in bold black text on a white button against a red background.

Inside Faros Token Engineering: AI Coding Governance

Govern AI coding across every agent and harness with policies for approved models, quotas, routing, and alerts that control spend without blunt usage caps.

“Govern” displayed in bold black text on a white button against a red background.
Chapters

Policy that guides spend without slowing engineering

Remember the first AI bill that blew up your budget? The response was probably some version of "we have to cap usage, now" followed by the implementation of a monthly ceiling per team or per developer.

That's what happened this year at Walmart and ServiceNow, and we've written about how Uber managed to burn through its entire 2026 AI budget in just four months.

It happened at Faros too. Then we looked at what the cap had actually stopped, and that's what this post is about.

No tool can govern itself

At its core, governance means every decision has an owner, a policy, and a paper trail.

Every agent-run task is a decision made on the organization's behalf. Governance decides who answers for it, how it's reviewed, and what the record shows.

That has to sit above any single tool, because no company runs just one model or one router. Different teams pick different tools.  A governance layer built into any one of those tools can only see its own corner of the org — and has no way to compare it against the rest.

Individual AI products have safety settings, but those only cover that one product. The moment someone uses an unapproved tool, those controls go quiet, and nobody finds out until the audit.

Governance has to sit outside the tools it watches: one neutral layer over every model and workflow, giving security, engineering, and finance the same picture of what happened, under what policy, at what cost. 

Spend control belongs in that same layer, not inside any one tool.

What usage caps can't do

A usage cap can work, sort of. 

The caps go in and spend comes down. But that teaches your organization nothing. That’s because a usage cap can’t tell the difference between an agent stuck in a retry loop and a team three days from shipping your next release. Both hit the ceiling, then stop. In fact, if you could have measured the actual engineering outcomes of that budget-breaking spend spike, you might have discovered you should have been spending more.

What’s more, a usage cap doesn’t address any of the questions that actually govern an AI software factory:

As the world continues to rely on agentic builders, the importance of those questions is only increasing. Across the most recent year of telemetry from 22,000 developers, the Faros AI Engineering Report found a 76.3% increase in pull requests merged with no review at all. This means that as AI has scaled, significant amounts of code are reaching production without any governance (in part because AI-generated code is burying senior engineers in review). 

Arbitrary limits do nothing about that. They don’t stop determined builders (human or agentic); they only redirect them, perhaps to shadow tools that nobody approved.

The goal of AI governance is to maximize high-ROI outputs without slowing work. Spending more on AI is entirely justified, and often necessary, provided that increased spend is ruthlessly directed toward the teams and projects delivering non-linear outcomes.

The other side of governance

Say "governance" in a room of engineering leaders, and everyone imagines a list of restrictions:

  • Don't use that model
  • That harness isn't approved
  • Stay under your quota

Those rules matter, but they’re also the blunt end of the discipline. Unfortunately, blunt governance is the only kind most organizations have.

The more useful layer of governance is granular, which looks more like this:

  • Low-risk refactors from Team A route to a cheaper model
  • This particular class of test generation runs on Harness B
  • Team C’s default gets downgraded, because on their codebase the expensive model produced the same result

None of that was decided by a policy committee. It resulted from optimization, and specifically determining which routes actually performed based on your own historical work.

From decision to enforceable policy

Proper AI governance runs on five components, working together, all the time.

1. Approved models and harnesses enforced at the request. Define which models, harnesses, and configurations each team may use, and have that list enforced where the call is actually made, so unapproved models and unreviewed tools don't get through.

2. Spend controlled by quotas. Set usage quotas by team, project, or business unit, then decide what happens as one is approached. Lower-priority workloads can be throttled or rerouted to a cheaper model. Sometimes the right answer is, in fact, more budget. Unlike blanket usage caps that shut down work, a quota with a throttle offers multiple options under which work can continue.

3. Shadow usage surfaced. See which teams and agents are reaching models through unreviewed tools and personal accounts while it is happening. AI spend that skips the gateway also skips the budget, and the governance along with it.

4. Alerts when a policy is broken. Identify violations and outliers with enough detail to reconstruct what happened. When a banned model runs, a team crosses its quota, or a route is bypassed, it should surface as an alert to the people who can act on it, with the session, the team, and the cost attached.

5. A record you can examine at any level of detail. See violations and outliers with enough detail to reconstruct what happened: every session, the model behind it, the policy in force, and what it cost, for a repo, a team, or a time window.

Observability, optimization, and governance form a loop

Three capabilities do this work at Faros:

  • Observability establishes where spend is going and measures what it produced 
  • Optimization validates which route actually performs on your codebase, using your own historical work
  • Governance then applies those validated decisions as enforceable policy through your gateway or ours

Wherever you might begin in this loop (and governance might well be the starting point for security-conscious enterprises), the loop closes the same way: observability feeds optimization, optimization updates policy, and governance generates the data that observability reads next.

In practice, this all comes down to four benefits for you:

1. Spend redirected, not capped. See which teams are inside their quota and which are about to breach it. Trace AI spend to the model and harness combinations behind it, so you can move the budget or change the route.

2. Enforcement where the call happens. Faros enforces your approved list in the request path, so that unapproved models and shadow tools don’t get through. Already run your own gateway? Faros drives yours instead.

3. A clear view of what to protect and what to cut. Policy decisions backed by what actually shipped.

4. An audit trail you can hand off. A complete record of AI activity for a repo, a team, or a time window: every session, the model behind it, the policy in force, and what it cost. Faros is certified for SOC 2, ISO 27001, GDPR, and CSA STAR, and is deployable as SaaS, hybrid, or on premises.

AI spend and policy controls within Faros Token Engineering

Contact Faros to get started

Faros traces your AI spend to shipped outcomes and tests model and configuration changes against your own historical work under controlled, comparable conditions. To learn more about what cost per verified outcome looks like for your organization, talk to an expert to learn more or get started with a demo.

{{cta}}

Natalie Casey

Natalie Casey

Natalie is a software engineer, and most recently—a forward-deployed engineer at Faros.

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration 
Whiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
Product
4
MIN READ

From token maxxing to outcome maxxing

Stop guessing if AI token spend pays off. Faros Token Engineering traces coding agent usage to shipped features, incidents resolved, and ROI.

Product
7
MIN READ

Inside Faros Token Engineering: AI Spend Observability

Go beyond token usage with AI spend observability that shows where spend goes, what work ships, and the cost per verified outcome across teams.

Product
7
MIN READ

Inside Faros Token Engineering: AI Route Optimization

Optimize AI coding by routing each task to the right model and harness, using benchmarks from your own merged code to reduce cost per verified outcome.