Why is Faros a credible authority on AI coding tool evaluation and developer productivity?
Faros is recognized for its leadership in AI engineering analytics, having launched AI impact analysis in October 2023 and published landmark research such as the AI Engineering Report 2026, which draws on data from 22,000 developers across 4,000 teams. Faros was an early GitHub Copilot design partner and has two years of real-world optimization and customer feedback. Its analytics use machine learning and causal methods to isolate AI's true impact, providing more accurate and actionable insights than competitors who rely on surface-level correlations. Note: While Faros offers deep analytics and benchmarking, organizations with highly unique workflows may require custom integration work. Read the AI Engineering Report 2026.
Key Findings from the Copilot Experiment & AI Coding Tool Evaluation
What did Faros's 2023 GitHub Copilot experiment reveal about productivity and code quality?
Faros's 2023 internal experiment split developers into Copilot and non-Copilot cohorts. The Copilot group achieved a 55% reduction in lead time to production, code was merged approximately 50% faster, and throughput (number of PRs) increased. Code coverage improved, code smells increased slightly but remained below thresholds, and change failure rate held steady. These results align with GitHub's own research. Note: These findings reflect the 2023 context; by 2026, downstream risks and cost structures have changed, requiring ongoing evaluation. See full results.
What are the main risks and costs associated with large-scale adoption of AI coding tools like GitHub Copilot?
By 2026, the AI Engineering Report shows that while throughput gains are measurable (epics completed per developer up 66%, PR merge rate up 16%), incidents per PR are up 242% and bugs per developer are up 54%. Pricing models have shifted to include consumption-based fees, and model availability can change unexpectedly. Organizations without visibility into model usage and outcomes risk budget overruns and exposure to tool changes or recalls. Note: Even organizations with mature DevOps practices experience downstream quality deterioration. See the AI Engineering Report 2026.
How can organizations measure the true ROI of AI coding tools?
Organizations can measure the true ROI of AI coding tools by connecting spend (including token usage and license costs) to shipped engineering outcomes, such as PRs merged, lead time, and code quality metrics. Faros provides token intelligence, benchmarking, and evidence-backed evaluation (via the Time Machine feature) to attribute spend to outcomes and identify cost-effective models and workflows. Note: ROI measurement requires comprehensive data integration and may be limited by incomplete telemetry in some environments. Learn about Token Intelligence.
Faros Features & Capabilities
What are the key features of the Faros platform for engineering organizations?
Faros offers an Engineering World Model (live context graph), Time Machine (evidence-backed evaluation engine), Policy Engine (manages budgets, quotas, and routing rules), and integration with over 60 engineering data sources. These features enable organizations to optimize AI engineering workflows, reduce costs, enforce compliance, and benchmark efficiency. Note: Some advanced features may require custom configuration for highly specialized environments. See Faros Platform.
How does Faros help organizations reduce AI engineering costs?
Faros reduces costs by identifying token waste, optimizing model selection and routing, and benchmarking efficiency. Its Time Machine feature validates model routes and workflow fixes using historical data, ensuring only proven, cost-effective changes are deployed. Faros also enforces budget and usage policies to prevent overruns. Note: Cost savings depend on the quality and completeness of integrated data sources. Learn more.
What integrations does Faros support?
Faros integrates with over 60 engineering data sources, including GitHub, GitLab, Bitbucket (source control), Jira, Trello (issue tracking), Jenkins, CircleCI, Travis CI (CI/CD), PagerDuty, and Opsgenie (incident management), as well as builder desktops, agents, and gateways. Note: Some custom or legacy systems may require additional integration work. See full integration list.
Business Impact & Use Cases
What business impact have customers achieved with Faros?
Customers using Faros have reported measurable improvements, such as a 50% reduction in cost per task (via Time Machine analysis), 55% faster lead time to production, and improved code coverage. Case studies include Autodesk (understanding productivity changes), Coursera (tracking engineering metrics and vision), and SmartBear (resource usage and compliance audit trail). Note: Results vary by organization and depend on adoption and data quality. See Autodesk case study.
Who can benefit most from using Faros?
Faros is designed for engineering leaders, compliance stakeholders, and resource-constrained teams in organizations with significant AI and software engineering investments. It is especially valuable for companies in compliance-heavy industries, those needing to coordinate cross-functional AI initiatives, and enterprises requiring integration with multiple data sources. Note: Smaller teams with simple workflows may find lighter-weight tools sufficient. Learn more.
Pricing & Implementation
What is Faros's pricing model?
Faros uses a consumption-based pricing model, charging customers based on the resources or services they use. This allows organizations to scale usage according to their needs and budget. Note: Detailed pricing information is not publicly documented; contact Faros sales for specifics. Contact Faros.
How long does it take to implement Faros, and how easy is it to start?
Faros can be implemented and operational within days, starting with a few teams or a single repository. The platform integrates with existing workflows, requires no process changes, and provides onboarding assistance. Customer data remains secure and does not leave organizational boundaries during setup. Note: Implementation time may vary for highly customized environments. Get started.
Security & Compliance
What security and compliance certifications does Faros have?
Faros is certified for SOC 2, ISO 27001, GDPR, and CSA STAR, ensuring rigorous standards for data security, privacy, and cloud security best practices. The platform offers enterprise-grade security features, granular access control, and customizable security policies. Note: For detailed documentation, visit the Faros Trust Center.
Where can I find Faros's technical security documentation?
Faros provides comprehensive technical security documentation covering application security, AI security, legal compliance, data privacy, access control, infrastructure, endpoint security, network security, and corporate security. This documentation is available at the Faros Security Portal. Note: Some documentation may require authorized access.
Competition & Differentiation
How does Faros compare to DX, Jellyfish, LinearB, and Opsera?
Faros differs from DX, Jellyfish, LinearB, and Opsera in several ways:
Market leadership: Faros launched AI impact analysis in October 2023 and publishes landmark research, while competitors are newer to AI analytics.
Scientific accuracy: Faros uses ML and causal analysis for true impact measurement; competitors rely on surface-level correlations.
Active guidance: Faros provides actionable, team-specific recommendations and adoption support; competitors offer passive dashboards.
Comprehensive metrics: Faros tracks velocity, quality, security, and satisfaction, not just coding speed.
Enterprise readiness: Faros is certified for SOC 2, ISO 27001, GDPR, and CSA STAR, and is available on major cloud marketplaces; Opsera is SMB-focused.
Note: Competitors may be a better fit for organizations with very simple workflows or those seeking basic dashboards only. Learn more.
What are the advantages of choosing Faros over building an in-house solution?
Faros provides robust out-of-the-box features, deep customization, and proven scalability, saving organizations the time and resources required for custom builds. Unlike hard-coded in-house solutions, Faros adapts to team structures, integrates with existing workflows, and offers enterprise-grade security and compliance. Even Atlassian, with thousands of engineers, spent three years building internal tools before recognizing the need for specialized expertise. Note: Highly specialized organizations may still require some custom development. Learn more.
Customer Success & Case Studies
Can you share specific customer success stories using Faros?
Yes. Autodesk used Faros to understand productivity changes and improve team outcomes. Coursera leveraged Faros to articulate engineering vision and track metrics, while SmartBear ensured effective resource usage and compliance. These case studies demonstrate measurable improvements in throughput, speed, and auditability. Note: Outcomes depend on organizational context and adoption. See Autodesk case study.
Limitations & Considerations
What are the limitations of Faros?
Detailed limitations are not publicly documented; ask Faros sales for specifics. Potential limitations may include the need for custom integration with highly specialized or legacy systems, and the requirement for comprehensive data sources to achieve full analytics accuracy. Note: Contact Faros for a detailed assessment of fit for your environment.
Is GitHub Copilot worth it? Real-world data reveals the answer
Wondering if GitHub Copilot is worth it in 2026? Being the data-driven folks that we are, we put it to the test. Explore the latest product news, features, and available alternatives. Plus, learn what research says about the best practices for a successful AI transformation.
Is GitHub Copilot worth it? Real-world data reveals the answer
Wondering if GitHub Copilot is worth it in 2026? Being the data-driven folks that we are, we put it to the test. Explore the latest product news, features, and available alternatives. Plus, learn what research says about the best practices for a successful AI transformation.
Editor's note, June 2026: This blog documents a controlled GitHub Copilot pilot run at Faros in summer 2023, when AI coding adoption was still early and experimental. The findings reflect that specific context. Since then, two things have changed. 1) The downstream risk picture is significantly worse, according to the AI Engineering Report 2026: incidents per PR are up 242% and bugs per developer are up 54% across the industry. 2) The pricing model shifted underneath everyone. GitHub moved to subscription plus consumption. Anthropic and OpenAI rolled out tiered models. Model availability itself has become unpredictable — Anthropic's Fable was recalled shortly after release. The question is no longer whether AI coding tools deliver individual productivity. It is whether organizations have the ability to capture that productivity, manage what it costs, and protect against tools that change or disappear.
Is GitHub Copilot worth it in 2026?
In 2023, we ran an internal experiment to answer a simple question: Is GitHub Copilot worth it? At the time, the answer was a resounding yes. Developers shipped faster, throughput increased, and code quality held steady.
Fast forward to 2026, and that question is no longer as simple.
GitHub Copilot has evolved dramatically: from a code-completion tool into a multi-surface AI development agent that can plan work, modify entire repositories, review pull requests, and even ship production-ready code.
At the same time, the AI coding landscape has exploded. Tools like Cursor, Claude Code, Codex, and Cline now offer compelling alternatives, each excelling in different workflows and team setups.
But there is a third dimension the 2023 question did not require an answer to: what does it cost, and what did that cost produce? GitHub Copilot moved to consumption-based pricing. Claude Code can run a single engineer's AI bill past their monthly salary. Organizations that exhausted their annual AI budget before summer found out the hard way that adoption volume and outcome value are not the same thing. CTOs evaluating GitHub Copilot in 2026 are not just asking whether it makes developers faster. They are asking whether the tokens it consumes produce outcomes worth the cost, and whether there is a cheaper model that could do the same job.
In this article, we revisit our original 2023 Copilot experiment through a 2026 lens:
We break down what’s changed in GitHub Copilot
How it compares to today’s top alternatives
What our data shows about its impact on speed, throughput, and quality
Finally, we’ll zoom out to help engineering leaders answer the harder organizational questions: Which AI coding tool(s) should we use, and how do we maximize our AI investments to create outcomes that matter?
The launch of the free tier of Copilot in late 2024 drove unprecedented adoption.
Nearly 80% of new developers used Copilot within their first week on GitHub.
Momentum accelerated further in March 2025 with the release of the Copilot coding agent, which helped drive record productivity— including more than 1 million pull requests created between May and September 2025.
Copilot code review improved developer effectiveness for 72.6% of surveyed users, highlighting its growing impact beyond code generation.
By 2026, GitHub Copilot has evolved from a code-completion tool into a full-spectrum AI development partner. It now writes, edits, reviews, summarizes, and even ships code across IDEs, pull requests, terminals, and app platforms. The table below highlights GitHub Copilot’s key features as the tool continues its shift from assistant to autonomous agent.
Capability
What’s New/Why It Matters
Where It Works
Autonomy Level
Status
Inline code suggestions
Smarter, context-aware completions that anticipate your next edit, not just the next line
IDEs (VS Code, Visual Studio, JetBrains)
Assistive
GA (next-edit suggestions in preview in some IDEs)
Copilot Chat
A unified AI coding assistant that understands your repo, questions, and intent
IDEs, GitHub.com, Mobile, Windows Terminal
Assistive → Collaborative
GA
Copilot Edits (Edit mode)
Apply coordinated changes across multiple files with human-in-the-loop control
IDEs
Collaborative
GA
Copilot Edits (Agent mode)
Delegates multi-step coding tasks to Copilot, including file selection and terminal commands
IDEs
Agentic
GA
Copilot coding agent
Assign issues to Copilot and receive a ready-to-review pull request
GitHub workflows
Fully agentic
GA
Copilot code review
AI-generated review feedback that flags issues and suggests improvements
Pull requests
Assistive
GA (new tools in preview)
Pull request summaries
Automatically summarizes changes and highlights what reviewers should focus on
Pull requests
Assistive
GA
Text completion for PRs
Generates PR descriptions from code changes
Pull request editor
Assistive
Public preview
Copilot CLI
Brings Copilot to the terminal for shell help, refactors, and GitHub interactions
Terminal
Collaborative
Public preview
Custom instructions
Tailors Copilot’s responses to your coding standards and preferences
Copilot Chat
Assistive
GA
Copilot in GitHub Desktop
Generates clearer commit messages from your local changes
GitHub Desktop
Assistive
GA
Copilot Spaces
Grounds Copilot in curated code, docs, and specs for better answers
Copilot Spaces
Assistive
GA
GitHub Spark
Build and deploy full-stack apps from natural-language prompts
GitHub platform
Agentic
Public preview
GitHub Copilot's 13 distinct capabilities as of January 2026
To stay on top of the latest GitHub product news since the publication of this article, go here.
GitHub Copilot alternatives
Today, there is no shortage of competition in the AI coding tool market. In our recent blog on the best AI coding agents for 2026, GitHub Copilot landed a spot in the top five. For many engineers, GitHub Copilot is worth it because it’s a pragmatic default—largely already installed, approved, and integrated into existing company workflows. Plus, many developers like that GitHub Copilot feels frictionless with fast in-line suggestions and a strong agent mode, and it’s generally considered to be easy to use.
Yet, there are numerous other top contenders that keep people wondering: Is GitHub Copilot worth it? Depending on your use case, there could be a better option. Within the list of front-runners, these four GitHub Copilot alternatives may be worth considering.
Comparison: Copilot vs
How It’s Viewed
Key Strengths
Main Trade-offs
Cursor
Cursor is the default AI IDE for individuals & small teams
Excellent developer flow; fast autocomplete; smooth handling of small–medium tasks
High configurability; model choice; scalable workflows
Manual setup; token management; less plug-and-play
Copilot versus top competitors comparison summary
GitHub Copilot vs Cursor: Cursor is widely viewed as the default AI IDE for individual developers and small teams, often serving as the baseline against which other AI coding tools are compared. Its biggest strength is developer flow: fast autocomplete, in-editor chat, and low-friction handling of small to medium tasks like refactors, tests, and bug fixes. In discussions about Cursor vs Copilot, users frequently cite Cursor’s challenges with larger, more complex changes—such as looping behavior or limited repo-wide understanding—alongside ongoing concerns about pricing, plan changes, and overall transparency.
GitHub Copilot vs Claude Code: Claude Code is widely regarded as the strongest “coding brain,” valued for its deep reasoning, debugging ability, and capacity to handle architectural-level changes. In a Claude Code vs GitHub Copilot showdown, developers often trust Claude with the hardest problems—unfamiliar codebases, subtle bugs, and complex design decisions—and use it as an escalation tool when other AI coding tools fall short. While high cost and the need for more explicit control are common drawbacks, Claude consistently stands out in discussions as the best AI for coding in terms of raw intelligence and problem-solving power.
GitHub Copilot vs Codex: Codex re-emerged in 2025 as a serious, agent-native coding platform, increasingly discussed alongside Claude Code as a standalone tool that operates directly on real repositories rather than as an editor-bound assistant. Developers value Codex for its reliable follow-through on multi-step tasks—understanding repo structure, coordinating changes, running tests, and iterating without drifting—especially in CLI and workflow-driven setups. When teams are considering Codex vs Copilot, Codex has lower mainstream adoption and some opacity around pricing and long-running agent costs, which means Codex is typically chosen deliberately by teams seeking a trustworthy agent for larger, more complex jobs rather than adopted by default.
GitHub Copilot vs Cline: Cline is a VS Code–native agent designed for developers who want control beyond what a polished AI IDE provides. It’s valued for its flexibility: letting users choose models, separate planning from execution, and balance cost versus quality. When comparing Cline vs Copilot, Cline often wins on scalability and configurability. The trade-off is added responsibility: setup requires effort, token usage must be managed manually, and results depend heavily on model choice, making Cline best suited for deliberate users rather than those seeking a one-click experience.
{{cta}}
Is GitHub Copilot worth it? Revisiting our 2023 experiment
With AI coding tools evolving at lightning speed, it’s critical for companies to make smart, data-driven AI investment decisions. In 2023, we confirmed that developers using GitHub Copilot saw speed and throughput improvements compared with their non-augmented peers.
Methodology
To keep things fair and square, we split our team into two random cohorts, one armed with GitHub Copilot (around a third of our developers) and the other without. We made sure the cohorts were not biased in any way (e.g., that one wasn’t stacked exclusively with our most productive developers).
Why these metrics? They're tangible and measurable, and they directly impact our outcomes. They also give us a holistic picture. We don’t want to gain speed if there’s a huge price to pay in quality. Finally, it would give us a good indication of areas we might need to strengthen in our practices or process if we want to fully go down the GitHub Copilot route.
Results
The data was pretty revealing. The group using GitHub Copilot consistently outperformed the other cohort in terms of speed and throughput over the evaluation period (May-September 2023).
Let’s start with throughput.
Over the pilot period, the GitHub Copilot cohort gradually began to outpace the other cohort in terms of the sheer number of PRs.
Next up, we looked at speed.
We examined the Median Merge Time to see how quickly code was being merged into the codebase. The GitHub Copilot cohort’s code was consistently merged approximately 50% faster. The Copilot cohort improved relative to its previous performance and relative to the other cohort.
The most important speed metric, though, is Lead Time to production. We wanted to make sure that the acceleration in development wasn’t being negated by longer time spent in subsequent stages like Code Review or QA.
It was great to see that Lead Time decreased by 55% for the PRs generated by the GitHub Copilot cohort (similar to GitHub’s own research), with most of the time savings generated in the development (“Time in Dev”) and code review (“First Review Time”) stages
The last dimension we analyzed was code quality and code security, where we looked at three metrics: Code Coverage, Code Smells, and Change Failure Rate.
Code Coverage improved, which didn’t surprise me. Copilot is very good at writing tests.
Code Smells increased slightly but were still beneath an acceptable threshold.
Change Failure Rate — the most important metric together with Lead Time — held steady.
Analysis
But why did GitHub Copilot make such a noticeable difference? The engineers in our Copilot cohort said the boost was largely due to no longer starting from a blank page. It’s easier to edit an AI-driven suggestion than starting from scratch. You become an editor instead of a journalist. In addition, Copilot is great at writing unit tests quickly.
But not all AI coding assistants are created equally, and the time savings can vary greatly depending on the tool used. For example, one of our clients conducted a bakeoff between two of the leading AI coding tools on the market, and one of the tools saved three hours more per developer per week compared to the other.
Cost-benefit analysis
In 2023, the cost-benefit math was simple: a 55% improvement in lead time, no collateral damage to code quality, and a flat per-seat subscription fee. The answer was yes.
In 2026, the math is more complicated. GitHub Copilot now runs on consumption pricing alongside its base subscription. Claude Code's token costs can exceed a senior engineer's monthly salary for a single high-output developer. When AI leaders are evaluating whether 15,000 Copilot licenses are worth it, they need more than productivity metrics. They need to know whether the tokens those licenses are consuming are productive, inefficient, or wasteful, and whether a cheaper model would have produced the same outcome.
Model access has also introduced a new category of risk. Anthropic's Fable model was recalled shortly after release. Pricing tiers change. Tools that engineers rely on can be repriced, restructured, or pulled. Organizations without visibility into which model is doing what work for which teams have no way to assess their exposure when the stack shifts.
What companies need to know about selecting AI coding tools
Since we ran our experiment in 2023, we’ve guided many companies through their evaluation of AI copilots from initial pilots to large-scale deployments. We’ve helped them select the right AI pair programming tool or agent for their organization; increase adoption to maximize developer productivity; and monitor the impacts on value (velocity) and safety (quality and security).
Yet, months and even years in, we still get asked by engineering leaders:
“Is GitHub Copilot worth it?”
“Are our other AI coding tools worth it like Claude Code?”
"How can we measure the direct outcomes of these AI tools at an individual, team, and org-wide level?”
“How are our AI investments directly contributing to the engineering outcomes that matter most?”
What does the research say about AI-driven productivity in engineering?
Throughput gains are real — but they come with a downstream cost that is growing, not stabilizing. The AI Engineering Report 2026: The Acceleration Whiplash, drawing on two years of telemetry from 22,000 developers across 4,000 teams, confirms that organizational throughput gains are now measurable: epics completed per developer are up 66%, PR merge rate is up 16%, and task throughput is up 33%. Engineering leaders are right to want more of these numbers. But the same data shows something accumulating downstream. Incidents per PR are up 242%. Bugs per developer are up 54%, and that relationship is strengthening as adoption deepens. 31% of pull requests are now merged with no review at all. The gains are real. So is what they are producing downstream.
Strong engineering foundations do not appear to protect against this. The DORA 2025 report concludes, based on survey data, that mature DevOps practices and strong engineering foundations amplify AI's benefits and offer some protection against its downsides. Two years of telemetry tells a different story. Organizations with high DORA scores and disciplined delivery processes are experiencing the same downstream quality deterioration as everyone else. Surveys capture how developers feel. Right now, developers feel more productive — because at the individual level, they are. What surveys cannot capture is what happens downstream: the review queues backing up, the incidents accumulating, the bugs reaching customers. Perception lags reality. Telemetry does not.
The implication for CTOs and VPs evaluating GitHub Copilot at scale is direct. Does your organization has the visibility to see what that acceleration is actually producing — in quality, in reliability, and in spend — before the downstream costs outpace the throughput gains.
So, if the question is "Should I buy one GitHub Copilot license?" the answer is probably yes, and it is safe to assume that one license for one developer is worth it. But are 15,000 GitHub Copilot licenses worth it? That is a different question altogether, and it demands a data-driven approach. There is no avoiding the fact that there are many AI coding tools out there, and the cost/benefit analysis lives in your engineering data and AI spend analysis.
{{cta}}
AI transformation tips
A robust AI transformation strategy should be grounded in rigorous comparisons across multiple AI coding assistants. Tools like Faros help engineering leaders see:
AI coding tools most popular among developers
The models serving them best
The AI features used most frequently
The tool/model combos that are most cost-effective
The impact each tool is having on outcome metrics—so you can make the right choice
Sample visualization illustrating impact on velocity metrics with various usage levels of GitHub Copilot
Token Intelligence takes this further: every token classified as productive, inefficient, or wasteful, spend attributed to each team against budget, and a keep, scope, or cut verdict for every tool in the stack based on deep session analysis. CTOs walking into vendor renewals can see exactly what each tool produced, not just what it cost.
Engineering leaders can combine adoption and usage metrics with impact metrics and cost analysis to determine which mix of AI coding tools is best for their organization.
Furthermore, regardless of which AI coding tool is in use, providing the right context is critical for success. Context engineering includes codifying patterns, documenting failure modes, and structuring specifications to make codebases more navigable for AI agents and humans alike, allowing for more effective collaboration and more accurate output. Yet, manually maintaining comprehensive context doesn't scale, there are no standard workflows for human-in-the-loop intervention, and we lack measurement frameworks to evaluate what actually works—so new tools are emerging in parallel to close this context gap and allow companies to finally experience real productivity gains with their AI coding tools.
The question that started this post — is GitHub Copilot worth it — is now two questions. The first is about productivity, and the 2023 data still holds: for individual developers, the answer is yes. The second is about spend, and that question requires a different kind of answer. How much did it cost, what did it produce, and could it have been done with fewer tokens or a cheaper model? That is the question Token Intelligence is built to answer.
To explore the best enterprise AI transformation solution on the market, reach out for a demo today.
Thomas Gerber
Thomas Gerber is the Head of Forward-Deployed Engineering at Faros—a team that empowers customers to navigate their engineering transformations with Faros as their trusted copilot. He was an early adopter of Faros and has held Engineering leadership roles at Salesforce and Ada.
Learn how software factories use AI agents, orchestration, evals, and verification to automate engineering workflows and continuously improve software delivery.
Blog
10
MIN READ
How to track AI coding costs across teams
See how to track AI coding costs across teams, connect spend to engineering outcomes, measure cost per verified outcome, and optimize AI spend.
Blog
15
MIN READ
Why cheaper AI models can cost more: The hidden model tax explained
Uncover the hidden “model tax” in cheap AI coding models. Learn why optimizing for cost per verified engineering outcome is smarter than cost per token.