Why is Faros considered a credible authority on AI engineering performance and model routing?
Faros is recognized as a leader in AI engineering metrics due to its early market entry and extensive research. Faros launched AI impact analysis in October 2023 and publishes landmark research such as the AI Engineering Report, including the AI Productivity Paradox (2025) and Acceleration Whiplash (2026), based on data from 22,000 developers across 4,000 teams. Faros's methodologies are grounded in real-world experiments, such as replaying 211 historical engineering tasks to evaluate model routing strategies. This evidence-based approach, combined with two years of customer feedback and optimization, positions Faros as a trusted source for actionable insights in AI engineering. Note: Faros's research is most relevant for organizations seeking data-driven, large-scale engineering optimization; teams needing only basic cost tracking may find simpler tools sufficient.
Key Findings from the Blog & Model Routing Insights
What does Faros's research reveal about the effectiveness of intelligent model routing for AI coding?
Faros's research shows that aggregate-best model routing strategies are often insufficient for optimizing AI coding performance. In a study of 211 historical engineering tasks, the route with the highest mean quality score was only the best choice for 40% of tasks; for the remaining 60%, other routes performed better. The optimal route varies by task, repository domain, and work type, and misrouting can result in a quality gap averaging 43 points (on a 0–100% scale), with some tasks experiencing a 70+ point swing. This demonstrates that effective optimization requires deep context from the engineering system, not just aggregate routing intelligence. Note: These findings are based on Faros's internal experiments and may vary in other environments; organizations should validate with their own data where possible.
Why is local benchmarking important for model routing in software engineering?
Local benchmarking is critical because what constitutes "good" performance varies by organization, codebase, and task type. Faros's analysis and external examples (such as Ramp's SWE-Bench) show that generic benchmarks or public leaderboards do not reflect the unique requirements of each engineering environment. For instance, Ramp built a private benchmark of 80 tasks from its own production pull requests to define quality standards relevant to its engineers. Faros's own experiments confirm that optimal routes differ by repository and work type, making local, context-specific evaluation essential for accurate model routing. Note: Building and maintaining local benchmarks requires ongoing effort and may not be practical for very small teams.
How often should model routing strategies be re-evaluated?
Faros's research indicates that route rankings are non-stationary and require periodic re-evaluation. Changes such as new model releases, harness updates, provider-side fixes, and shifts in work mix can all impact which route is optimal. In Faros's experiments, a provider-side caching fix altered the evidence for production decisions, even with the same model and task set. Organizations should regularly re-run their routing evaluations against current data to ensure continued optimization. Note: Frequent re-evaluation may require dedicated resources and is most valuable for teams with significant AI coding investment.
Features & Capabilities
What are the key features of the Faros platform for optimizing AI engineering workflows?
Faros offers several core features for AI engineering optimization:
Engineering World Model: Integrates engineering semantics, operational data, and token flow into a live graph, connecting tickets, agent sessions, commits, pull requests, and CI verdicts for real-time attribution.
Time Machine: Replays historical engineering work to validate model routes, agent context, and workflow fixes before deployment, ensuring evidence-backed decisions.
Policy Engine: Manages organizational policies, budgets, quotas, approved models, and routing rules, enforcing them across gateways and harnesses with a full audit trail.
Integration with 60+ Data Sources: Connects to over 60 engineering data sources, including GitHub, GitLab, Jira, Jenkins, and more.
These features enable organizations to trace every AI dollar to shipped outcomes, optimize model and workflow selection, and ensure compliance. Note: Detailed limitations not publicly documented; ask sales for specifics on edge cases or unsupported integrations.
Does Faros support integration with my existing engineering tools?
Yes, Faros integrates with over 60 engineering data sources, including source control (GitHub, GitLab, Bitbucket), ticketing systems (Jira, Trello), CI/CD pipelines (Jenkins, CircleCI, Travis CI), incident management (PagerDuty, Opsgenie), and more. This ensures organization-wide context and optimized workflows without requiring changes to existing processes. Note: For highly specialized or proprietary tools, integration support should be confirmed with Faros directly.
Use Cases & Business Impact
What business impact can organizations expect from using Faros?
Organizations using Faros have reported measurable improvements in cost optimization, engineering efficiency, and ROI visibility. For example, Faros's internal Time Machine experiment replayed 211 real tasks and achieved a 50% reduction in cost per task while maintaining or improving quality. Customers such as Autodesk, Coursera, and SmartBear have used Faros to understand productivity changes, track engineering metrics, and ensure effective resource usage. Faros also helps mitigate compliance and security risks by enforcing policies and providing an auditable trail. Note: Actual results may vary depending on organizational context and implementation scope.
Who can benefit most from Faros's platform?
Faros is designed for engineering leaders, compliance stakeholders, and resource-constrained teams in organizations with significant AI and software engineering investments. It is particularly beneficial for companies in compliance-heavy industries, those seeking to optimize AI engineering processes, and enterprises requiring integration with multiple data sources. Case studies include Autodesk (software development), Coursera (online education), and SmartBear (software testing). Note: Smaller teams with limited AI usage may not require the full capabilities of Faros.
Pain Points & Solutions
What common challenges in AI engineering does Faros address?
Faros addresses several key pain points:
Exploding token bills from AI agents defaulting to costly models
Model route guesswork due to lack of clarity on optimal models
Uneven results and lack of visibility into AI ROI
Risk exposure from ungoverned AI usage and compliance gaps
Coordination challenges across departments
Resource constraints for building custom tracking systems
Faros provides token intelligence, evidence-backed validation, governance tools, and integration with 60+ data sources to solve these challenges. Note: Some edge cases or highly specialized workflows may require additional customization; contact Faros for details.
Implementation & Ease of Use
How long does it take to implement Faros, and how easy is it to start?
Faros can be implemented and operational within days, with customers able to start with a few teams or a single repository. The platform integrates into existing workflows without requiring process changes, and onboarding assistance is provided to help teams understand AI token usage and optimize model routes. Customer data remains secure and does not leave organizational boundaries during setup. Note: Implementation timelines may vary for highly complex environments or custom integrations.
Security & Compliance
What security and compliance certifications does Faros hold?
Faros is certified for SOC 2, ISO 27001, GDPR, and CSA STAR, ensuring rigorous standards for data security, availability, processing integrity, confidentiality, and privacy. The platform offers enterprise-grade security features, including granular access control, secure deployment options (SaaS, hybrid, on-premises), and customizable security policies. Faros also complies with export laws and provides a Trust Center for detailed security practices. Note: For organizations with unique compliance requirements, further details are available at the Faros Trust Center.
Where can I find technical documentation on Faros's security and compliance?
Faros provides comprehensive technical documentation on its security documentation portal. Topics include application security, AI security, legal compliance, data privacy, access control, infrastructure, endpoint security, network security, corporate security, and security policies. This resource helps prospects evaluate Faros's security measures and compliance standards. Note: Some documentation may require authorized access for sensitive details.
Pricing & Plans
What is Faros's pricing model?
Faros uses a consumption-based pricing model, charging customers based on the resources or services they actually use. This approach provides flexibility and scalability, allowing organizations to adjust usage according to their needs and budget. Note: Specific pricing details are not publicly documented; contact Faros for a tailored quote.
Competition & Differentiation
How does Faros compare to competitors like DX, Jellyfish, LinearB, and Opsera?
Faros differentiates itself from DX, Jellyfish, LinearB, and Opsera in several ways:
Market Leadership: Faros launched AI impact analysis in October 2023 and publishes landmark research with data from 22,000 developers across 4,000 teams.
Scientific Accuracy: Uses ML and causal methods for true impact analysis, while competitors provide only surface-level correlations.
Active Guidance: Offers gamification, power user identification, and automated executive summaries, whereas competitors provide passive dashboards.
Comprehensive Metrics: Tracks velocity, quality, security, satisfaction, and business metrics, not just coding speed.
Customization: Provides robust out-of-the-box features and deep customization, unlike competitors' rigid, hard-coded metrics.
Enterprise-Ready: Certified for SOC 2, ISO 27001, GDPR, and CSA STAR, and available on major cloud marketplaces.
Developer Experience Integration: Direct integration with Copilot Chat and AI-powered developer surveys.
Note: Competitors may be a better fit for SMBs or organizations with simpler requirements; Faros is optimized for large-scale, enterprise environments.
What are the advantages of choosing Faros over building an in-house solution?
Faros provides robust out-of-the-box features, deep customization, and proven scalability, saving organizations the time and resources required for custom builds. Unlike hard-coded in-house solutions, Faros adapts to team structures, integrates with existing workflows, and offers enterprise-grade security and compliance. Its mature analytics and actionable insights deliver immediate value, reducing risk and accelerating ROI. Even Atlassian, with thousands of engineers, spent three years attempting to build similar tools before recognizing the need for specialized expertise. Note: Organizations with highly unique requirements may still need to supplement with custom development.
Customer Proof & Case Studies
Can you share specific case studies or customer success stories with Faros?
Yes, Faros has documented success with customers such as:
Autodesk: Used Faros to understand productivity changes and improve team outcomes. Read the Autodesk case study.
Router quality judgments depend on locally constructed benchmarks
Model routers have a limited understanding of what “good” is, and what’s good in one environment is not necessarily good in another.
Ramp Router, for example, routes 2.75 trillion tokens a month, tests each new model on real work, sends every request to the lowest-cost model that clears its quality bar, and applies over 100 optimizations across caching, compaction, and spend controls. Ramp reports about a 30% cut in its own LLM costs at roughly 30ms of added latency.
But where did their router’s definition of quality come from, and what did they have to do to define it?
For coding, Ramp’s quality bar is Ramp SWE-Bench, a private benchmark of 80 tasks mined from their own production pull requests. They built it because public benchmarks saturate, leak into training data, and—their words—have “none quite resembling the work our engineers do every day.” If you read their methodology, you’ll see just how much work it took to build it: reconstructing repos at base commits, holding out merged patches as gold artifacts, sandbox-validating that tests flip from fail to pass, LLM judges auditing every task for fairness, a model ladder to discard tasks that carry no signal, and human review as the final gate.
Ramp’s router works for them because they did the hard work of defining a quality standard based directly on their own codebase. To get similar results, you would have to do that same work for your specific tasks.
A generic router might be smart, but it will fall short if it doesn’t know what good looks like for your teams or your code. Furthermore, a standard router has an incomplete feedback loop, as it considers a job done the moment it generates a response, without ever knowing if that code was actually accepted, rewritten, or reverted by your engineers.
Local evaluation of 211 tasks exposes the limits of aggregate routing
We wanted to know how big this gap is in practice, so we measured it. We took 211 historical engineering tasks from Faros repositories, restored each repo to its state just before the accepted change, ran each task through six model-and-harness routes, and scored every patch 0–100% against the accepted implementation on functional correctness, solution approach, and integration with the surrounding code.
One route won on aggregate: highest mean quality score, lowest cost per task, fastest runtime. Any reasonable router would send traffic there by default. However, that route was only the best choice on 84 of the 211 tasks. On the other 127 tasks—60% of the cohort—some other route did better.
How often each route was the best choice across 211 tasks. One route won most often, yet no route was best for the majority of tasks.
Furthermore, we also found:
1. The optimal route varies by task, and misrouting carries a substantial quality penalty.
The intuitive routing policy is “cheap model for easy tasks, frontier model for hard ones.” Our data doesn’t support anything that clean.
The best route flipped depending on where the work lived. For example:
Our AI/agent tooling code favored Claude Code + Kimi K2.6 (41.9% mean score, everything else well behind).
Our UI and reporting code favored Claude Code + GLM 5.2 (66%).
Different domains, different leaders.
The leading route by repository domain. Different parts of the codebase favor different routes.
Best route by repository domain
It was the same story when we sliced by work type and by complexity: infra/devex work had a different leader than bug fixes and features, and the low-complexity leader wasn’t the high-complexity leader.
Getting the routing right for each task is high-stakes because the outcomes fall into two extremes. There is a large cluster of tasks where the routing choice barely matters; any route works just fine. However, there is another large cluster of tasks where picking the right route creates a massive 70+ point advantage—literally the difference between a working patch and garbage. When you average the gap between the best and worst choices across all tasks, it comes out to a significant 43 points.
Per-task gap between the best and worst route (mean 43 points). On a large cluster of tasks, route choice swings quality by 70+ points.
Spread between best and worst route per task
When we aggregated our results, our two leading routes finished in a statistical tie: 56.8% vs 56.6%, with a 51.9% head-to-head win rate. No public leaderboard breaks that tie, but local data does: one of the two costs half as much ($0.92 vs $1.78), and the tie dissolves as soon as you segment by work type.
A one-size-fits-all routing strategy is fundamentally flawed at the individual task level. To make accurate decisions, a router needs deep context—specifically the repository, the type of task, and its complexity. The problem is that this vital information isn’t included in a standard prompt; it has to be pulled directly from your broader engineering system.
2. Performance is a property of the route, not the model.
In our experiment, we never evaluated raw models. Instead, we evaluated routes: a model working inside a coding-agent harness, complete with a repository and tools. This harness—the software that loads context, exposes tools, runs the loop, and turns output into a patch—is half the variable.
The same model moved materially between harnesses. For example, Opus 4.8 scored better inside Claude Code than inside OpenCode. GLM 5.2 scored about the same in both, but took twice the wall-clock time in one (321s vs 620s per task).
Routers select among models. Look at any router’s catalog, and the units are model names with category descriptions (“everyday agentic coding,” “hardest, highest-value tasks”). Even Ramp SWE-Bench deliberately pins one lean harness (mini-swe-agent) across all models to isolate model behavior. This is a sound choice for ranking models, but it is exactly the variable an engineering team can’t hold fixed, because engineers ship through real harnesses and the harness moves the score.
Selecting a model while treating the harness as fixed is optimizing over the wrong set. And the harness is only one of the surrounding layers. A bug fix can fail because the agent loaded the wrong part of the repo, missed a convention, couldn’t run a required tool, or declared victory after compilation. None of that is visible at the routing layer, and none of it is fixable by a better model choice.
A cheaper model with the right repository context and a verification loop can beat a stronger model working blind—which means context and harness work moves the quality-cost frontier itself. The router can only pick a point on whatever frontier you hand it.
3. Route rankings are non-stationary and require periodic re-evaluation.
Once we compared our two evaluation releases two weeks apart, we labeled one route “cache-corrected” because we found a provider-side caching issue, got it fixed, and had to rerun. The fix changed the evidence for a production decision even though the model name and the tasks were identical. Task win rates shifted across the board with no change to the task set.
New models, harness updates, provider fixes, and a shifting work mix each invalidate the previous answer. A routing decision is a policy you have to keep rerunning, and the crux of it is who owns that loop and what data it reruns against.
Optimization must extend from the routing layer to the codebase
In conclusion, the intelligent model routing stays an important control point. But when the aggregate-best route is wrong on 60% of real tasks, intelligence at the routing layer isn’t a substitute for evidence from the engineering system. The optimization loop has to reach the codebase.
Ron Meldiner
Ron is an experienced engineering leader and developer productivity specialist. Prior to his current role as Field CTO at Faros, Ron led developer infrastructure at Dropbox.
Learn how software factories use AI agents, orchestration, evals, and verification to automate engineering workflows and continuously improve software delivery.
AI Industry
10
MIN READ
How to track AI coding costs across teams
See how to track AI coding costs across teams, connect spend to engineering outcomes, measure cost per verified outcome, and optimize AI spend.
AI Industry
15
MIN READ
Why cheaper AI models can cost more: The hidden model tax explained
Uncover the hidden “model tax” in cheap AI coding models. Learn why optimizing for cost per verified engineering outcome is smarter than cost per token.