Policy that guides spend without slowing engineering
Remember the first AI bill that blew up your budget? The response was probably some version of "we have to cap usage, now" followed by the implementation of a monthly ceiling per team or per developer.
That's what happened this year at Walmart and ServiceNow, and we've written about how Uber managed to burn through its entire 2026 AI budget in just four months.
It happened at Faros too. Then we looked at what the cap had actually stopped, and that's what this post is about.
No tool can govern itself
At its core, governance means every decision has an owner, a policy, and a paper trail.
Every agent-run task is a decision made on the organization's behalf. Governance decides who answers for it, how it's reviewed, and what the record shows.
That has to sit above any single tool, because no company runs just one model or one router. Different teams pick different tools. A governance layer built into any one of those tools can only see its own corner of the org — and has no way to compare it against the rest.
Individual AI products have safety settings, but those only cover that one product. The moment someone uses an unapproved tool, those controls go quiet, and nobody finds out until the audit.
Governance has to sit outside the tools it watches: one neutral layer over every model and workflow, giving security, engineering, and finance the same picture of what happened, under what policy, at what cost.
Spend control belongs in that same layer, not inside any one tool.
What usage caps can't do
A usage cap can work, sort of.
The caps go in and spend comes down. But that teaches your organization nothing. That’s because a usage cap can’t tell the difference between an agent stuck in a retry loop and a team three days from shipping your next release. Both hit the ceiling, then stop. In fact, if you could have measured the actual engineering outcomes of that budget-breaking spend spike, you might have discovered you should have been spending more.
What’s more, a usage cap doesn’t address any of the questions that actually govern an AI software factory:
- Which model and harness combination was used?
- Was that the optimal combination for that task?
- Did the work produce a measurable outcome?
- Was the tool it ran through ever reviewed by security?
- Under which policy did any of the work happen?
As the world continues to rely on agentic builders, the importance of those questions is only increasing. Across the most recent year of telemetry from 22,000 developers, the Faros AI Engineering Report found a 76.3% increase in pull requests merged with no review at all. This means that as AI has scaled, significant amounts of code are reaching production without any governance (in part because AI-generated code is burying senior engineers in review).
Arbitrary limits do nothing about that. They don’t stop determined builders (human or agentic); they only redirect them, perhaps to shadow tools that nobody approved.
The goal of AI governance is to maximize high-ROI outputs without slowing work. Spending more on AI is entirely justified, and often necessary, provided that increased spend is ruthlessly directed toward the teams and projects delivering non-linear outcomes.
The other side of governance
Say "governance" in a room of engineering leaders, and everyone imagines a list of restrictions:
- Don't use that model
- That harness isn't approved
- Stay under your quota
Those rules matter, but they’re also the blunt end of the discipline. Unfortunately, blunt governance is the only kind most organizations have.
The more useful layer of governance is granular, which looks more like this:
- Low-risk refactors from Team A route to a cheaper model
- This particular class of test generation runs on Harness B
- Team C’s default gets downgraded, because on their codebase the expensive model produced the same result
None of that was decided by a policy committee. It resulted from optimization, and specifically determining which routes actually performed based on your own historical work.
From decision to enforceable policy
Proper AI governance runs on five components, working together, all the time.
1. Approved models and harnesses enforced at the request. Define which models, harnesses, and configurations each team may use, and have that list enforced where the call is actually made, so unapproved models and unreviewed tools don't get through.
2. Spend controlled by quotas. Set usage quotas by team, project, or business unit, then decide what happens as one is approached. Lower-priority workloads can be throttled or rerouted to a cheaper model. Sometimes the right answer is, in fact, more budget. Unlike blanket usage caps that shut down work, a quota with a throttle offers multiple options under which work can continue.
3. Shadow usage surfaced. See which teams and agents are reaching models through unreviewed tools and personal accounts while it is happening. AI spend that skips the gateway also skips the budget, and the governance along with it.
4. Alerts when a policy is broken. Identify violations and outliers with enough detail to reconstruct what happened. When a banned model runs, a team crosses its quota, or a route is bypassed, it should surface as an alert to the people who can act on it, with the session, the team, and the cost attached.
5. A record you can examine at any level of detail. See violations and outliers with enough detail to reconstruct what happened: every session, the model behind it, the policy in force, and what it cost, for a repo, a team, or a time window.
Observability, optimization, and governance form a loop
Three capabilities do this work at Faros:
- Observability establishes where spend is going and measures what it produced
- Optimization validates which route actually performs on your codebase, using your own historical work
- Governance then applies those validated decisions as enforceable policy through your gateway or ours
Wherever you might begin in this loop (and governance might well be the starting point for security-conscious enterprises), the loop closes the same way: observability feeds optimization, optimization updates policy, and governance generates the data that observability reads next.
In practice, this all comes down to four benefits for you:
1. Spend redirected, not capped. See which teams are inside their quota and which are about to breach it. Trace AI spend to the model and harness combinations behind it, so you can move the budget or change the route.
2. Enforcement where the call happens. Faros enforces your approved list in the request path, so that unapproved models and shadow tools don’t get through. Already run your own gateway? Faros drives yours instead.
3. A clear view of what to protect and what to cut. Policy decisions backed by what actually shipped.
4. An audit trail you can hand off. A complete record of AI activity for a repo, a team, or a time window: every session, the model behind it, the policy in force, and what it cost. Faros is certified for SOC 2, ISO 27001, GDPR, and CSA STAR, and is deployable as SaaS, hybrid, or on premises.

Contact Faros to get started
Faros traces your AI spend to shipped outcomes and tests model and configuration changes against your own historical work under controlled, comparable conditions. To learn more about what cost per verified outcome looks like for your organization, talk to an expert to learn more or get started with a demo.
{{cta}}





