Frequently Asked Questions

Product Information & Authority

Why is Faros a credible authority on AI coding agent failures and engineering productivity?

Faros is recognized for its leadership in AI engineering analytics, having launched AI impact analysis in October 2023 and published landmark research such as the AI Engineering Report 2026, which draws on data from 22,000 developers across 4,000 teams. Faros's methodology is grounded in rigorous, rubric-based evaluation and causal analysis, enabling precise diagnosis of where AI coding agents fail and how to improve outcomes. The platform's insights are validated by real-world customer case studies from companies like Autodesk, Coursera, and SmartBear. Note: While Faros provides deep analytics and actionable recommendations, individual project results may vary based on data quality and organizational adoption. Read the AI Engineering Report 2026.

What is the main finding from Faros's analysis of AI coding agent failures?

Faros's analysis of approximately 4,000 agent failures across six models revealed that the most common failure mode is not a model capability gap, but rather instruction conflicts—specifically, agents misreading or over-literally applying task instructions. For example, a boilerplate instruction like "DO NOT MODIFY: Tests, configuration files" led agents to refuse required work, resulting in 247 failures. This insight highlights the importance of prompt engineering and harness design over simply upgrading models. Note: These findings are based on a specific dataset and harness; results may differ in other environments. Read the full analysis.

Features & Capabilities

How does Faros help diagnose and resolve AI coding agent failures?

Faros uses a rubric-based evaluation system that scores agent outputs against per-task requirements, providing detailed explanations for every failure. Its Time Machine feature replays historical engineering work to validate model routes, agent context, and workflow fixes before deployment. This approach enables organizations to identify root causes of failures—such as instruction conflicts or change hygiene issues—and implement targeted improvements. Note: The effectiveness of these diagnostics depends on the quality and completeness of historical engineering data available for analysis.

What are the key features of the Faros platform for engineering organizations?

Key features of Faros include the Engineering World Model (a live graph connecting tickets, agent sessions, commits, pull requests, and CI verdicts), the Time Machine (evidence-backed evaluation engine), and the Policy Engine (manages policies, budgets, quotas, and routing rules). Faros integrates with over 60 engineering data sources, provides cost optimization, efficiency benchmarking, and a unified control plane for observability, optimization, and governance. Note: Some advanced features may require integration with specific data sources or workflows. Learn more about Faros Platform.

How does Faros integrate with existing engineering tools and workflows?

Faros connects to over 60 engineering data sources, including source control (GitHub, GitLab, Bitbucket), ticketing systems (Jira, Trello), CI/CD pipelines (Jenkins, CircleCI, Travis CI), incident management (PagerDuty, Opsgenie), and builder desktops/agents. This broad integration ensures organization-wide context and enables Faros to operate without requiring workflow changes. Note: Integration depth may vary depending on the specific tool and configuration. See full integration list.

Use Cases & Business Impact

What business impact can engineering organizations expect from using Faros?

Organizations using Faros have reported measurable improvements such as faster shipping of production code, reduced token waste, and maximized engineering outcomes. For example, Faros's internal Time Machine analysis replayed 211 real tasks across seven model and harness routes, resulting in a 50% reduction in cost per task while maintaining or improving quality. Customers like Autodesk, Coursera, and SmartBear have used Faros to understand productivity changes, track metrics, and ensure effective resource usage. Note: Actual impact depends on organizational context and adoption. See case studies.

Who can benefit most from Faros's platform?

Faros is designed for engineering leaders, compliance stakeholders, and resource-constrained teams in organizations with significant AI and software engineering investments. It is especially valuable for companies in compliance-heavy industries, those needing to coordinate cross-functional AI initiatives, and enterprises requiring integration with multiple engineering data sources. Case studies highlight its use in software development (Autodesk), online education (Coursera), and software testing (SmartBear). Note: Teams with highly specialized or proprietary workflows may require additional customization.

What pain points does Faros address for engineering teams?

Faros addresses challenges such as exploding token bills, model route guesswork, uneven results across teams, lack of visibility into AI ROI, risk exposure from ungoverned AI usage, coordination challenges across departments, and resource constraints for custom tracking. Its features—like token intelligence, Time Machine, and governance tools—help organizations optimize spend, validate model choices, and ensure compliance. Note: Some pain points may require process changes or additional data integration for full resolution.

Competition & Comparison

How does Faros compare to DX, Jellyfish, LinearB, and Opsera?

Faros differs from DX, Jellyfish, LinearB, and Opsera in several ways: it was first to market with AI impact analysis (October 2023), publishes landmark research, and uses causal analysis for accurate ROI measurement. Faros provides end-to-end tracking (velocity, quality, security, satisfaction, business metrics), active adoption support, and enterprise-grade compliance (SOC 2, ISO 27001, GDPR, CSA STAR). Competitors often offer only surface-level correlations, limited integrations (mainly Jira and GitHub), and passive dashboards. Faros's out-of-the-box dashboards and flexible customization support complex enterprise needs. Note: Faros may require more initial configuration for highly customized environments; competitors may be simpler for SMBs with basic needs.

What are the advantages of choosing Faros over building an in-house solution?

Faros offers mature analytics, robust out-of-the-box features, deep customization, and proven scalability, saving organizations the time and resources required for custom builds. Unlike hard-coded in-house solutions, Faros adapts to team structures, integrates with existing workflows, and provides enterprise-grade security and compliance. Even Atlassian, with thousands of engineers, spent three years trying to build developer productivity tools in-house before recognizing the need for specialized expertise. Note: In-house solutions may be preferable for organizations with highly unique requirements not addressed by commercial platforms.

Technical Requirements & Implementation

How long does it take to implement Faros, and how easy is it to start?

Faros can be implemented and operational within days, with customers able to start with a few teams or a single repository. The platform integrates into existing workflows without requiring process changes, and onboarding assistance is provided to help teams understand AI token usage and optimize model routes. Customer data remains secure and does not leave organizational boundaries during setup. Note: Implementation time may vary for organizations with complex or highly customized environments. Get started with Faros.

What technical documentation and support resources are available for Faros?

Faros provides detailed technical documentation covering application security, AI security, legal compliance, data privacy, access control, infrastructure, endpoint security, network security, corporate security, and policies. The documentation is available at the Faros Security Portal. Onboarding assistance and customer support are also provided during implementation. Note: Some documentation may require access permissions. Visit the Faros Security Portal.

Security & Compliance

What security and compliance certifications does Faros hold?

Faros is compliant with SOC 2, ISO 27001, GDPR, and CSA STAR standards, ensuring rigorous data security, availability, processing integrity, confidentiality, and privacy. The platform offers enterprise-grade security features, granular access control, secure deployment options (SaaS, hybrid, on-premises), and customizable security policies. Faros also complies with export laws and provides a Trust Center for detailed security practices. Note: For the most current certification status, refer to the Faros Trust Center.

Pricing & Plans

What is Faros's pricing model?

Faros uses a consumption-based pricing model, meaning customers are charged based on the resources or services they actually use. This provides flexibility and scalability, allowing organizations to adjust usage according to their needs and budget. Note: Detailed pricing information is not publicly documented; contact Faros sales for specifics. Learn more about Faros.

Why AI coding agents actually fail (it's not the model)

Why do coding agents fail? We analyzed 4,000 errors across 6 models and discovered the real culprits.

magnifying glass hovering over X's in circles signifying investigating failures

Why AI coding agents actually fail (it's not the model)

Why do coding agents fail? We analyzed 4,000 errors across 6 models and discovered the real culprits.

magnifying glass hovering over X's in circles signifying investigating failures
Chapters

Where coding agents actually go wrong

We clustered ~4,000 agent failures across six models. The single biggest cause wasn't a capability gap. It was one sentence in the prompt.

TL;DR: Everyone quotes agent pass rates. Almost nobody can tell you why the failures happened, because a binary test verdict has nothing to say beyond "no." Our rubric-based judge does: every failed requirement comes with a specific reason. So we took six coding agents, ran each on the same 100 SWE-bench Pro tasks, and clustered every one of the ~4,000 failed rubric items by what was asked and why it failed. The headline finding was telling on where optimization often lies: the largest failure cluster across all six models is agents obeying a boilerplate instruction so literally that they refuse work the task explicitly requires. This won’t be fixed by a smarter model, but by reading your failures.

Why should you care?

If you've been following this series, you know the setup: real engineering work mostly can't be scored by running tests, so we score patches against per-task rubrics, which are weighted checklists of binary requirements, graded by a blinded LLM judge. We've shown that this graded score sees quality inside failures, and that when it disagrees with a benchmark's tests, the benchmark is sometimes the one that's wrong.

This post is about another valuable property of the same instrument. A test harness gives you one bit per task. A rubric gives you a verdict per requirement—did the right files change, was the spec satisfied, were tests weakened, does the behavior hold—and for every failed requirement, the judge writes down why. Multiply that across six models and 100 tasks and you get 1,858 scored requirements per model, 11,148 graded verdicts, and roughly 4,000 failures, each with an explanation attached.

That's not a leaderboard. That's a diagnosis. And when you cluster those explanations, the picture of "where agents go wrong" on your codebase becomes a gold mine.

The setup: six agents, 100 tasks, every failure explained

We ran six models (gemini-flash-lite-3.1, claude-haiku-4-5, gpt-5-mini, gemini-flash-3, gemini-pro-3.1, and claude-opus-4-7) on the same 100 SWE-bench Pro tasks under the same harness. Each patch was graded against the task's rubric, whose items fall into four categories:

  • FC: file-change: did the right files get touched, and only those?
  • SA: spec-alignment: does the change do what the PR asked?
  • I:  integrity: no weakened tests, no unrelated churn, no broken compatibility.
  • R: runtime: does the behavior actually hold?

Overall item success was 64.5%, ranging from 52.6% (gemini-flash-lite) to 75.3% (claude-opus-4-7). Those are the numbers a leaderboard would report. Everything below is what the leaderboard can't see.

Rubrics success rate, per model
Rubrics success rate, per model

Finding 1: The #1 failure mode is an instruction conflict, not a capability gap

The largest explanation cluster across all six models—247 failures—is agents misreading or over-literally applying task instructions, and one instruction dominates: the harness boilerplate that says "DO NOT MODIFY: Tests, configuration files."

That line exists for a good reason: it stops agents from deleting failing assertions to make their patch "pass." But many SWE-bench Pro rubrics—mirroring the actual merged PRs—require adding test cases or updating an assertion so a boundary test stays meaningful. The agent is now holding two contradictory instructions, and we can watch it choose, in its own words. Here's claude-haiku-4-5 on a Teleport task that required updating a payload-size boundary test:

Example of an agent taking instructions too literally

The agent then shipped a patch that left a boundary test exercising nothing. 

On other tasks, agents wrote perfectly good test cases - into throwaway scripts in /tmp, where the rubric can't see them and no CI ever will. On future-architect/vuls, this single misreading—"do not modify tests" interpreted as "do not add tests"—is the dominant failure mode for the entire repository, 52 failures on its own.

Another example of an agent failing to add tests

Note this cluster is not a weak-model problem. It's the #1 explanation theme for claude-haiku-4-5 (56 failures) and for claude-opus-4-7 (63 failures)—the best model in the study. If anything, the stronger model reasons its way into the refusal more articulately. You cannot buy your way out of a contradictory prompt with a bigger model. The fix costs one sentence of prompt engineering—distinguishing "don't weaken existing tests" from "don't touch anything test-shaped"—and in this dataset that sentence is worth more than some model upgrades.

Finding 2: Agents fail like sloppy engineers, not stupid ones

The second-largest cluster, 141 failures, is pure change hygiene: temporary debug scripts, helper files, and unrelated edits swept into the final diff—very often by a single reflexive git add -A. One agent fixed a NodeBB function correctly, then committed its scratch test file alongside the fix and introduced quote-style inconsistencies because it edited the source with sed.

Example of an issue with change hygiene

That sed detail generalizes. On the runtime axis, the top explanation cluster is unsafe automated text editing: brittle search-and-replace commands that target the wrong lines, overwrite whole files, or silently fail to apply - leaving half-modified implementations the agent never re-read. The code that results isn't wrong because the model couldn't reason about the problem. It's wrong because the edit never landed the way the agent believed it did, and nothing in its loop made it check.

Read enough of these and a pattern emerges: the failures look less like a junior engineer who doesn't understand the codebase and more like a rushed senior one who doesn't verify their own work—staging blindly, editing mechanically, submitting without a final read of the diff. That's actually good news. "Doesn't verify" is a harness problem with harness solutions: a mandatory git status review before submission, structured editing instead of sed, a self-check pass on the final diff. None of it requires a smarter model.

Finding 3: What separates strong models from weak ones (and what doesn't)

With failures broken out by rubric axis, you can see exactly where capability lives:

Failure rate by rubric axis × model (lower is better). Items per axis: FC 540, SA 498, I 372, R 448.

Three things jump out. Integrity is nearly a solved problem for everyone. Even the weakest model avoids weakening tests or breaking compatibility three times out of four; the spread across models is small. Agents have largely learned not to vandalize.

Spec-alignment and runtime correctness are where the money goes. These axes separate strong from weak models most sharply: flash-lite fails spec-alignment at more than twice Opus's rate. When you pay for a frontier model, what you're buying is follow-through: the change does what the PR actually asked, and the behavior actually holds.

File-change discipline is hard for everyone. Touching exactly the right files, and nothing else, is the highest failure rate on the board even for the best model, at 31.9%. Precision of scope, not raw problem-solving, is the frontier's remaining weakness.

Finding 4: Upgrading a model removes its quirks, not the hard tasks

Here's our favorite cut of the data. For every failed item, we counted how many other models also failed it. 297 requirements were failed by all six models, genuinely hard asks, from preserving an obscure end-of-role marker in Ansible's PlayIterator to handling a space-switch timing gap in Element's room list.

Now look at each model's failures through that lens. For gemini-flash-lite, 20.5% of failures are unique to it—items every other model handled. For claude-opus-4-7, that number is 2.2%. Ten failures out of 458. Meanwhile, roughly two-thirds of Opus's failures are the universal 297 that nobody solved.

Each model's failures split by how many other models also failed the item. 297 items failed by all 6 = genuinely hard; lightest = unique to that model.

That's a precise statement of what a model upgrade buys you: it makes the idiosyncratic failures disappear. What's left over is task hardness—under-specified requirements, deep repository context, genuinely subtle bugs—and no model on the market gets you past those. If your agent keeps failing a class of tasks, the first question isn't "which model is next"; it's "would any model pass this, or is the task the problem?" With per-item cross-model data, that stops being a philosophical question and becomes a lookup.

One more comparison drives it home. The spread between the best and worst model is 22.7 points of item success. The spread between the easiest and hardest repository in the study—qutebrowser at 77.6% versus tutao/tutanota at 35.5%—is 42.1 points. Where your codebase sits matters roughly twice as much as which frontier model you pick. Model selection isn't a global question; it's a per-repository one.

Success rate by models and repos

What we're not claiming

Honesty notes, in the tradition of this series.

The judge has a noise floor. 167 items (13%) split exactly 3/3 across the six models—the band most sensitive to judge variance—and near-boundary conclusions there deserve caution. The clusters we've highlighted are the largest ones, well clear of that band, but individual "why it failed" explanations are LLM judgments, not ground truth. We audit this judge continuously precisely because we make claims like these on top of it.

One harness, one run per model. All six models ran under the same scaffold, which is what makes the cross-model comparison clean. But it also means some failure modes (the one-command-at-a-time constraint that pushed agents toward sed, the boilerplate that created the test-file conflict) are properties of the harness-plus-model pair, not the model alone. That's not a weakness of the analysis; it's its point. Your agents also run inside a harness, and its fingerprints are all over their failures too.

Cluster themes are summaries, not measurements. The counts are exact; the prose descriptions of each cluster are model-written characterizations of its members, which we spot-checked against the underlying transcripts.

The takeaway: Stop counting failures and start reading them

Across ~4,000 failures, the recurring lesson is that a pass rate hides the one thing you can act on. "Opus scores 75%, Haiku scores 60%" tells you what a model tier costs. "Your harness prompt is vetoing required test work, your submission step is committing scratch files, and 297 of your tasks are unwinnable as written" tells you what to fix: this week, for free.

And every one of those findings came out of the same rubric scores we already produce for ranking. The explanations were sitting in the failure data; clustering them is cheap. If you're evaluating agents with any judge that produces reasons—and if it doesn't, that's worth fixing first—group your failures by what was asked and why it failed before you spend another dollar on a model comparison. You'll likely find, as we did, that your biggest "model problems" aren't.

This diagnostic layer is part of what Time Machine now produces. When we replay your engineering history to compare models and harnesses on your own repositories, the same clustering runs over your failures: which instructions your agents trip on, which repositories punish them, which tasks no model can pass as specified. You don't just learn which model wins on your work - you learn why the others lose, and how much of that is yours to fix. 

Contact us for a demo.

Thierry Donneau-Golencer

Thierry Donneau-Golencer

Thierry is Head of Product at Faros, where he builds solutions to empower teams and drive engineering excellence. His previous roles include AI research (Stanford Research Institute), an AI startup (Tempo AI, acquired by Salesforce), and large-scale business AI (Salesforce Einstein AI).

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration 
Whiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
AI Industry
12
MIN READ

What is a software factory? How it works

Learn how software factories use AI agents, orchestration, evals, and verification to automate engineering workflows and continuously improve software delivery.

AI Industry
10
MIN READ

How to track AI coding costs across teams

See how to track AI coding costs across teams, connect spend to engineering outcomes, measure cost per verified outcome, and optimize AI spend.

AI Industry
15
MIN READ

Why cheaper AI models can cost more: The hidden model tax explained

Uncover the hidden “model tax” in cheap AI coding models. Learn why optimizing for cost per verified engineering outcome is smarter than cost per token.