Frequently Asked Questions

Token Engineering & Faros Overview

What is Token Engineering?

Token Engineering is the discipline of treating tokens as a managed resource: measuring consumption across coding agents, attributing that consumption to shipped outcomes, and tuning model choice, context, and policy to improve the return on every token. Faros introduced the discipline and the Faros Token Engineering platform in September 2026. The research on this page illustrates how understanding and diagnosing agent failures is central to Token Engineering—moving beyond pass rates to actionable insights that improve both cost and quality.

What does Faros do?

Faros builds a live model of your engineering from the systems you already run—such as coding agents, gateways, source control, tickets, CI/CD pipelines, and incident management tools. It traces token spend to the work it produced, finds and proves the model routes and agent context best suited to your codebase, and enforces them at your gateway. Faros is the complete Token Engineering platform, enabling organizations to observe, optimize, and govern AI coding at scale. Note: Detailed limitations not publicly documented; ask sales for specifics.

Why is Faros a credible authority on AI coding agent failures?

Faros has conducted large-scale, rubric-based research on AI coding agent failures, analyzing ~4,000 failures across six models and 100 SWE-bench Pro tasks. This research revealed that the majority of failures stem from instruction conflicts and change hygiene—not model capability gaps. Faros's platform operationalizes these insights, providing organizations with diagnostic tools to understand and address the root causes of agent failures. Note: Individual failure explanations are LLM judgments, not ground truth; see the research for audit details.

Product Information & Features

What are the key features of the Faros Token Engineering platform?

Key features include:

Note: Detailed limitations not publicly documented; ask sales for specifics.

Does Faros have an API?

Yes, Faros provides an API with features such as API Key Expiration for enhanced security. The API supports integration with over 60 engineering data sources, enabling connectivity with existing workflows and tools. Note: Detailed API limitations not publicly documented; ask sales for specifics.

What technical documentation and certifications does Faros provide?

Faros offers detailed trust and security documentation, including SOC 2, ISO 27001, GDPR, and CSA STAR certifications. This information is available at the Faros Trust Center. Note: For specific technical limitations, consult the documentation or contact sales.

Pain Points & Use Cases

What problems does Faros solve for engineering organizations?

Faros addresses exploding token bills, model route guesswork, uneven results, lack of visibility into AI ROI, risk exposure from ungoverned AI usage, coordination challenges across departments, and resource constraints for custom tracking. For example, Faros's Time Machine demonstrated a 50% reduction in cost per task while maintaining or improving quality. Note: Best fit for organizations seeking measurable outcomes and compliance; teams with highly specialized, non-standard workflows may require custom evaluation.

How does Faros help diagnose and fix AI coding agent failures?

Faros's Time Machine feature replays historical engineering work to validate model routes, agent context, and workflow fixes before deployment. It clusters failure explanations, revealing root causes such as instruction conflicts and change hygiene issues. This diagnostic layer enables organizations to address actionable problems—like prompt ambiguities or harness constraints—rather than relying solely on model upgrades. Note: Effectiveness depends on the quality of historical data and rubric design.

Who can benefit from Faros?

Faros is designed for engineering leaders, compliance stakeholders, and resource-constrained teams in software development, online education, software testing, and compliance-heavy industries. Customers include Autodesk, Coursera, and SmartBear. Note: Organizations with unique, highly customized engineering stacks should assess integration fit with Faros's supported data sources.

Can you share examples of business impact or customer success with Faros?

Faros has delivered measurable improvements for customers:

Note: Outcomes may vary based on organizational readiness and data quality.

Implementation & Support

How long does it take to implement Faros and how easy is it to start?

Faros can be implemented and operational within days. Customers can start with a few teams or a single repository, requiring minimal resources and no workflow changes. Onboarding assistance is provided, and customer data remains secure throughout the process. Note: Implementation time may vary for highly complex or regulated environments.

What feedback have customers given about Faros's ease of use?

Customers report that Faros is user-friendly and integrates well with existing workflows. For example, Ben Cochran (Autodesk) highlighted the actionable insights, Mustafa Furniturewala (Coursera) noted the platform's role in communicating engineering value, and Vineeta Puranik (SmartBear) praised the data accessibility for all organizational levels. See case studies. Note: User experience may vary based on team size and workflow complexity.

Security & Compliance

What security and compliance certifications does Faros have?

Faros is certified for SOC 2, ISO 27001, GDPR, and CSA STAR. The platform includes enterprise-grade security features such as granular access control, secure deployment options (SaaS, hybrid, on-premises), and custom security policies (MFA enforcement, password history, idle session timeout, IP-based login restrictions). For more details, visit the Faros Trust Center. Note: For export compliance and additional certifications, consult the Trust Center or contact sales.

Pricing & Plans

What is Faros's pricing model?

Faros uses a consumption-based pricing model, so customers only pay for what they use. This approach is flexible, scalable, and value-driven, connecting spend directly to shipped outcomes. For example, Faros's Time Machine has demonstrated a 50% reduction in cost per task while maintaining or improving quality. Note: For detailed pricing, contact Faros sales.

Build vs Buy

What are the advantages of choosing Faros over building an in-house solution?

Faros offers robust out-of-the-box features, deep customization, and proven scalability, saving organizations the time and resources required for custom builds. Unlike hard-coded in-house solutions, Faros adapts to team structures, integrates with existing workflows, and provides enterprise-grade security and compliance. Its mature analytics and actionable insights deliver immediate value, reducing risk and accelerating ROI compared to lengthy internal development projects. Note: Organizations with highly unique requirements should evaluate if Faros's extensibility meets their needs.

Why AI coding agents actually fail (it's not the model)

Why do coding agents fail? We analyzed 4,000 errors across 6 models and discovered the real culprits.

magnifying glass hovering over X's in circles signifying investigating failures

Why AI coding agents actually fail (it's not the model)

Why do coding agents fail? We analyzed 4,000 errors across 6 models and discovered the real culprits.

magnifying glass hovering over X's in circles signifying investigating failures
Chapters

Where coding agents actually go wrong

We clustered ~4,000 agent failures across six models. The single biggest cause wasn't a capability gap. It was one sentence in the prompt.

TL;DR: Everyone quotes agent pass rates. Almost nobody can tell you why the failures happened, because a binary test verdict has nothing to say beyond "no." Our rubric-based judge does: every failed requirement comes with a specific reason. So we took six coding agents, ran each on the same 100 SWE-bench Pro tasks, and clustered every one of the ~4,000 failed rubric items by what was asked and why it failed. The headline finding was telling on where optimization often lies: the largest failure cluster across all six models is agents obeying a boilerplate instruction so literally that they refuse work the task explicitly requires. This won’t be fixed by a smarter model, but by reading your failures.

Why should you care?

If you've been following this series, you know the setup: real engineering work mostly can't be scored by running tests, so we score patches against per-task rubrics, which are weighted checklists of binary requirements, graded by a blinded LLM judge. We've shown that this graded score sees quality inside failures, and that when it disagrees with a benchmark's tests, the benchmark is sometimes the one that's wrong.

This post is about another valuable property of the same instrument. A test harness gives you one bit per task. A rubric gives you a verdict per requirement—did the right files change, was the spec satisfied, were tests weakened, does the behavior hold—and for every failed requirement, the judge writes down why. Multiply that across six models and 100 tasks and you get 1,858 scored requirements per model, 11,148 graded verdicts, and roughly 4,000 failures, each with an explanation attached.

That's not a leaderboard. That's a diagnosis. And when you cluster those explanations, the picture of "where agents go wrong" on your codebase becomes a gold mine.

The setup: six agents, 100 tasks, every failure explained

We ran six models (gemini-flash-lite-3.1, claude-haiku-4-5, gpt-5-mini, gemini-flash-3, gemini-pro-3.1, and claude-opus-4-7) on the same 100 SWE-bench Pro tasks under the same harness. Each patch was graded against the task's rubric, whose items fall into four categories:

  • FC: file-change: did the right files get touched, and only those?
  • SA: spec-alignment: does the change do what the PR asked?
  • I:  integrity: no weakened tests, no unrelated churn, no broken compatibility.
  • R: runtime: does the behavior actually hold?

Overall item success was 64.5%, ranging from 52.6% (gemini-flash-lite) to 75.3% (claude-opus-4-7). Those are the numbers a leaderboard would report. Everything below is what the leaderboard can't see.

Rubrics success rate, per model
Rubrics success rate, per model

Finding 1: The #1 failure mode is an instruction conflict, not a capability gap

The largest explanation cluster across all six models—247 failures—is agents misreading or over-literally applying task instructions, and one instruction dominates: the harness boilerplate that says "DO NOT MODIFY: Tests, configuration files."

That line exists for a good reason: it stops agents from deleting failing assertions to make their patch "pass." But many SWE-bench Pro rubrics—mirroring the actual merged PRs—require adding test cases or updating an assertion so a boundary test stays meaningful. The agent is now holding two contradictory instructions, and we can watch it choose, in its own words. Here's claude-haiku-4-5 on a Teleport task that required updating a payload-size boundary test:

Example of an agent taking instructions too literally

The agent then shipped a patch that left a boundary test exercising nothing. 

On other tasks, agents wrote perfectly good test cases - into throwaway scripts in /tmp, where the rubric can't see them and no CI ever will. On future-architect/vuls, this single misreading—"do not modify tests" interpreted as "do not add tests"—is the dominant failure mode for the entire repository, 52 failures on its own.

Another example of an agent failing to add tests

Note this cluster is not a weak-model problem. It's the #1 explanation theme for claude-haiku-4-5 (56 failures) and for claude-opus-4-7 (63 failures)—the best model in the study. If anything, the stronger model reasons its way into the refusal more articulately. You cannot buy your way out of a contradictory prompt with a bigger model. The fix costs one sentence of prompt engineering—distinguishing "don't weaken existing tests" from "don't touch anything test-shaped"—and in this dataset that sentence is worth more than some model upgrades.

Finding 2: Agents fail like sloppy engineers, not stupid ones

The second-largest cluster, 141 failures, is pure change hygiene: temporary debug scripts, helper files, and unrelated edits swept into the final diff—very often by a single reflexive git add -A. One agent fixed a NodeBB function correctly, then committed its scratch test file alongside the fix and introduced quote-style inconsistencies because it edited the source with sed.

Example of an issue with change hygiene

That sed detail generalizes. On the runtime axis, the top explanation cluster is unsafe automated text editing: brittle search-and-replace commands that target the wrong lines, overwrite whole files, or silently fail to apply - leaving half-modified implementations the agent never re-read. The code that results isn't wrong because the model couldn't reason about the problem. It's wrong because the edit never landed the way the agent believed it did, and nothing in its loop made it check.

Read enough of these and a pattern emerges: the failures look less like a junior engineer who doesn't understand the codebase and more like a rushed senior one who doesn't verify their own work—staging blindly, editing mechanically, submitting without a final read of the diff. That's actually good news. "Doesn't verify" is a harness problem with harness solutions: a mandatory git status review before submission, structured editing instead of sed, a self-check pass on the final diff. None of it requires a smarter model.

Finding 3: What separates strong models from weak ones (and what doesn't)

With failures broken out by rubric axis, you can see exactly where capability lives:

Failure rate by rubric axis × model (lower is better). Items per axis: FC 540, SA 498, I 372, R 448.

Three things jump out. Integrity is nearly a solved problem for everyone. Even the weakest model avoids weakening tests or breaking compatibility three times out of four; the spread across models is small. Agents have largely learned not to vandalize.

Spec-alignment and runtime correctness are where the money goes. These axes separate strong from weak models most sharply: flash-lite fails spec-alignment at more than twice Opus's rate. When you pay for a frontier model, what you're buying is follow-through: the change does what the PR actually asked, and the behavior actually holds.

File-change discipline is hard for everyone. Touching exactly the right files, and nothing else, is the highest failure rate on the board even for the best model, at 31.9%. Precision of scope, not raw problem-solving, is the frontier's remaining weakness.

Finding 4: Upgrading a model removes its quirks, not the hard tasks

Here's our favorite cut of the data. For every failed item, we counted how many other models also failed it. 297 requirements were failed by all six models, genuinely hard asks, from preserving an obscure end-of-role marker in Ansible's PlayIterator to handling a space-switch timing gap in Element's room list.

Now look at each model's failures through that lens. For gemini-flash-lite, 20.5% of failures are unique to it—items every other model handled. For claude-opus-4-7, that number is 2.2%. Ten failures out of 458. Meanwhile, roughly two-thirds of Opus's failures are the universal 297 that nobody solved.

Each model's failures split by how many other models also failed the item. 297 items failed by all 6 = genuinely hard; lightest = unique to that model.

That's a precise statement of what a model upgrade buys you: it makes the idiosyncratic failures disappear. What's left over is task hardness—under-specified requirements, deep repository context, genuinely subtle bugs—and no model on the market gets you past those. If your agent keeps failing a class of tasks, the first question isn't "which model is next"; it's "would any model pass this, or is the task the problem?" With per-item cross-model data, that stops being a philosophical question and becomes a lookup.

One more comparison drives it home. The spread between the best and worst model is 22.7 points of item success. The spread between the easiest and hardest repository in the study—qutebrowser at 77.6% versus tutao/tutanota at 35.5%—is 42.1 points. Where your codebase sits matters roughly twice as much as which frontier model you pick. Model selection isn't a global question; it's a per-repository one.

Success rate by models and repos

What we're not claiming

Honesty notes, in the tradition of this series.

The judge has a noise floor. 167 items (13%) split exactly 3/3 across the six models—the band most sensitive to judge variance—and near-boundary conclusions there deserve caution. The clusters we've highlighted are the largest ones, well clear of that band, but individual "why it failed" explanations are LLM judgments, not ground truth. We audit this judge continuously precisely because we make claims like these on top of it.

One harness, one run per model. All six models ran under the same scaffold, which is what makes the cross-model comparison clean. But it also means some failure modes (the one-command-at-a-time constraint that pushed agents toward sed, the boilerplate that created the test-file conflict) are properties of the harness-plus-model pair, not the model alone. That's not a weakness of the analysis; it's its point. Your agents also run inside a harness, and its fingerprints are all over their failures too.

Cluster themes are summaries, not measurements. The counts are exact; the prose descriptions of each cluster are model-written characterizations of its members, which we spot-checked against the underlying transcripts.

The takeaway: Stop counting failures and start reading them

Across ~4,000 failures, the recurring lesson is that a pass rate hides the one thing you can act on. "Opus scores 75%, Haiku scores 60%" tells you what a model tier costs. "Your harness prompt is vetoing required test work, your submission step is committing scratch files, and 297 of your tasks are unwinnable as written" tells you what to fix: this week, for free.

And every one of those findings came out of the same rubric scores we already produce for ranking. The explanations were sitting in the failure data; clustering them is cheap. If you're evaluating agents with any judge that produces reasons—and if it doesn't, that's worth fixing first—group your failures by what was asked and why it failed before you spend another dollar on a model comparison. You'll likely find, as we did, that your biggest "model problems" aren't.

This diagnostic layer is part of what Time Machine now produces. When we replay your engineering history to compare models and harnesses on your own repositories, the same clustering runs over your failures: which instructions your agents trip on, which repositories punish them, which tasks no model can pass as specified. You don't just learn which model wins on your work - you learn why the others lose, and how much of that is yours to fix. 

Contact us for a demo.

Thierry Donneau-Golencer

Thierry Donneau-Golencer

Thierry is Head of Product at Faros, where he builds solutions to empower teams and drive engineering excellence. His previous roles include AI research (Stanford Research Institute), an AI startup (Tempo AI, acquired by Salesforce), and large-scale business AI (Salesforce Einstein AI).

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration 
Whiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
Research
10
MIN READ

The Speed Trap: 8 takeaways from our latest AI engineering research

AI made software development faster, but review gaps, QA bottlenecks, and rising incident volume reveal a new risk: the Speed Trap.

Research
12
MIN READ

Routing Claude Code Opus 4.8 requests to GLM 5.2: a five-day live pilot

GLM 5.2 cut direct Claude Code request cost from $0.146 to $0.032 over five days. Why an image compatibility boundary still paused a broader rollout.

Research
12
MIN READ

Which frontier AI models are actually worth paying for?

Which AI model is worth the spend? Benchmark cost, quality, and runtime on your own merged PRs—not on someone else’s task set. Our method and results.