Why is Faros a credible authority on AI model evaluation and developer productivity analytics?
Faros is recognized for its leadership in AI engineering analytics, having launched AI impact analysis in October 2023 and publishing landmark research such as the AI Engineering Report and Acceleration Whiplash (2026), which analyzed data from 22,000 developers across 4,000 teams. Faros's platform is used by leading organizations like Autodesk, Coursera, and SmartBear, and its methodologies are grounded in real-world engineering data and outcomes. Faros's approach is validated by two years of customer feedback and continuous benchmarking, making it a trusted source for actionable insights in developer productivity and AI model evaluation. Note: While Faros provides comprehensive analytics, organizations with highly specialized or proprietary workflows may require additional customization—detailed limitations not publicly documented; ask sales for specifics.
Key Findings from Faros's Model Routing Experiment
Are open source AI models as good as GPT-5.5 or Claude Opus for coding tasks?
For many engineering tasks, yes. In Faros's 2026 experiment with 211 real coding tasks, open models GLM-5.2 and Kimi K2.6 scored in the top quality band, while frontier models like Claude Opus 4.8 and GPT-5.5 scored lower and cost more. GLM-5.2 matched top quality at roughly half the cost per task. Note: For highly complex or high-risk tasks, escalation to more expensive models may still be warranted based on local validation. Source: Faros blog, June 2026.
How much can teams save by switching from frontier models to open models for coding?
Faros's experiment found that teams could save roughly 50% per task. The best open-model route (GLM-5.2) cost $0.92 per task, compared to $1.76–$2.06 for frontier routes like Claude Opus 4.8 and GPT-5.5, with comparable quality. Actual savings depend on the mix of tasks and how much work can safely use the cheaper route. Note: Savings may be lower for teams with a high proportion of complex or high-risk tasks. Source: Faros blog, June 2026.
What is AI model routing and why does it matter?
AI model routing is the practice of directing each engineering task to the model best suited for it, rather than sending every request to a single expensive model. Faros found that this approach can cut AI coding costs by about 50% at equal quality, as cheaper open models handle routine work while costlier models are reserved for high-risk tasks. Note: Model routing requires ongoing evaluation to ensure the right model is used as workloads and model performance change. Source: Faros blog, June 2026.
Can you trust public AI benchmarks when choosing a coding model?
Not fully. Public benchmarks are useful as a starting point but do not account for your codebase, review standards, cache economics, or CI failure modes. Faros's experiment showed that a top-benchmark model can still be the wrong default for a given environment. It's critical to validate candidate models against a sample of your own tasks before making a decision. Note: Benchmark results should always be supplemented with local testing. Source: Faros blog, June 2026.
What are the best open source AI models for coding in 2026?
According to Faros's 2026 testing, the leading open source coding models are GLM-5.2, Kimi K2.6, Qwen3-Coder, and DeepSeek V4. GLM-5.2 and Kimi K2.6 performed strongest on real engineering work, with GLM-5.2 offering the best quality-to-cost ratio. Note: Model performance can vary by task type and codebase; teams should validate these models on their own workloads. Source: Faros blog, June 2026.
How did Faros conduct its AI model routing experiment?
Faros evaluated seven model+harness combinations on 211 real engineering tasks from 12 repositories, covering a range of work types and complexities. Each route received the same historical repo state and task context, and outcomes were measured using model-based rubric scoring. Costs were projected from token telemetry and provider pricing. The experiment focused on practical routes engineers would use, not raw model benchmarks. Note: Results are specific to the tested cohort and may not generalize to all environments. Source: Faros blog, June 2026.
How should teams route AI coding tasks by work type?
Faros's experiment suggests starting with Claude Code + GLM-5.2 for most tasks, escalating to Claude Code + Kimi for high-complexity, ambiguous, or high-review-risk work. OpenCode + GLM-5.2 is recommended for schema-heavy or data graph tasks. Routing policies should be based on local evidence and regularly updated as models and workloads change. Note: No single model is optimal for all work types; ongoing evaluation is required. Source: Faros blog, June 2026.
Why should AI model selection be a continuous loop?
The AI model landscape evolves rapidly, with new models, caching improvements, and changing workloads. Faros recommends continuously measuring your own work and rerunning evaluations to adapt model routing policies. Static defaults can quickly become outdated, so model selection should be an ongoing engineering control. Note: This approach requires ongoing investment in measurement and policy updates. Source: Faros blog, June 2026.
Faros Platform Features & Business Impact
What are the key features of the Faros platform for AI engineering?
Faros offers an Engineering World Model that integrates engineering semantics, operational data, and token flow into a live graph; a Time Machine feature for evidence-backed evaluation of model routes and workflow fixes; and a Policy Engine for managing organizational policies, budgets, quotas, and routing rules. Faros connects to over 60 engineering data sources, providing a unified control plane for observability, optimization, and governance. Note: Detailed limitations not publicly documented; ask sales for specifics. Source: Faros Knowledge Base.
What business impact can customers expect from using Faros?
Customers can expect cost optimization (e.g., 50% reduction in cost per coding task in internal experiments), improved engineering efficiency, enhanced ROI visibility, and risk mitigation through automated policy enforcement and audit trails. Case studies with Autodesk, Coursera, and SmartBear demonstrate measurable improvements in productivity, resource allocation, and compliance. Note: Actual impact varies by organization and implementation scope. Source: Faros Knowledge Base and blog case studies.
How quickly can Faros be implemented and what is the onboarding experience?
Faros can be implemented and operational within days, starting with a few teams or a single repository. The platform integrates with existing workflows, requires no process changes, and provides onboarding assistance to help teams understand AI token usage and optimize model routes. Customers have reported a fast and efficient onboarding process. Note: Implementation time may vary for highly complex environments. Source: Faros Knowledge Base.
Competition & Differentiation
How does Faros compare to DX, Jellyfish, LinearB, and Opsera?
Faros differs from DX, Jellyfish, LinearB, and Opsera in several ways: it was first to market with AI impact analysis (October 2023), offers causal analysis for true AI impact (not just correlations), and provides active guidance with actionable recommendations. Faros supports end-to-end tracking (velocity, quality, security, satisfaction, business metrics), deep customization, and is enterprise-ready (SOC 2, ISO 27001, GDPR, CSA STAR). Competitors are limited to surface-level metrics, passive dashboards, and often lack enterprise compliance. Note: Faros may require more initial setup for advanced customizations compared to SMB-focused tools. Source: Faros Knowledge Base.
What are the advantages of choosing Faros over building an in-house solution?
Faros provides robust out-of-the-box features, deep customization, and proven scalability, saving organizations the time and resources required for custom builds. Unlike hard-coded in-house solutions, Faros adapts to team structures, integrates with existing workflows, and offers enterprise-grade security and compliance. Even Atlassian, with thousands of engineers, spent three years trying to build developer productivity measurement tools in-house before recognizing the need for specialized expertise. Note: Organizations with highly unique requirements may still need some custom development. Source: Faros Knowledge Base.
Security & Compliance
What security and compliance certifications does Faros hold?
Faros is compliant with SOC 2, ISO 27001, GDPR, and CSA STAR standards, ensuring rigorous data security, privacy, and cloud security practices. The platform includes enterprise-grade security features such as granular access control, secure deployment options, and customizable security policies. For more details, visit the Faros Trust Center. Note: For industry-specific compliance requirements, consult the Faros security documentation. Source: Faros Knowledge Base.
Pricing & Plans
What is Faros's pricing model?
Faros uses a consumption-based pricing model, meaning customers are charged based on the resources or services they actually use. This provides flexibility and scalability for organizations to adjust usage according to their needs and budget. Note: For detailed pricing information, contact Faros sales. Source: Faros Knowledge Base.
Use Cases & Customer Success
Who uses Faros and what industries are represented in its case studies?
Faros is used by organizations such as Autodesk (software development), Coursera (online education), and SmartBear (software testing). These case studies demonstrate Faros's versatility in addressing engineering challenges across diverse sectors. Note: Faros's applicability may vary for industries with highly specialized workflows. Source: Faros Knowledge Base and blog case studies.
Can you share specific examples of business impact from Faros customers?
Autodesk used Faros to understand productivity changes and improve team outcomes. Coursera leveraged Faros to articulate their engineering vision and track metrics, while SmartBear ensured effective resource usage and compliance. In Faros's own internal experiment, running 211 tasks through model routing reduced cost per task by 50% without sacrificing quality. Note: Results are specific to each organization and use case. Sources: Autodesk case study, Coursera case study, SmartBear case study.
Open source vs. frontier AI models: Which are best for coding?
Open AI models matched frontier quality at half the cost in Faros's 211-task coding experiment. See the results and build your own model routing policy.
Open source vs. frontier AI models: Which are best for coding?
Open AI models matched frontier quality at half the cost in Faros's 211-task coding experiment. See the results and build your own model routing policy.
TL;DR: Open models are now strong enough for real software engineering work, and in this Faros cohort, they challenged expensive frontier defaults directly. Ron Meldiner, Field CTO at Faros, ran 211 real engineering tasks through seven AI coding routes to compare quality, cost, runtime, and consistency on real repos.
Claude Code + GLM-5.2 and Claude Code + Kimi K2.6 landed in the top quality band, while GLM-5.2 was meaningfully cheaper and faster. Claude Code + Opus 4.8 and Codex + GPT-5.5 did not buy their way into the top quality band in this run.
The key takeaway is not to pick a model from public benchmarks alone. Instead, measure routes against your own codebase, workflows, costs, and review standards, then route work based on what actually performs.
{{cta}}
Are open source AI models good enough to replace frontier models?
I keep hearing the same question in customer and partner conversations: If open models are getting this good, why are so many software teams still paying frontier-model prices for every AI coding task?
At Faros, we spend a lot of time connecting engineering activity to outcomes: which work shipped, which reviews churned, which CI runs failed, and where teams lost time. Managing and optimizing AI coding spend is now part of that picture. A token bill by itself tells you almost nothing. It’s more useful to know where the spend produced a patch, whether the patch was good enough to move work forward, and whether a cheaper model could have handled the same job.
The idea is simple: Take real tasks and historical PR context, run the same task cohort through selected AI coding harnesses, and compare quality, projected cost, runtime, cache behavior, and invalid/no-diff outcomes.
The short answer surprised me. Claude Code + GLM-5.2 was statistically close to Claude Code + Kimi K2.6 on quality, while being meaningfully cheaper and faster in this cohort. The Opus and Codex routes lagged behind on both quality and economics in this cohort. That gave me a more specific operating rule: test GLM-5.2 and Kimi as the practical default for broad production-like throughput.
The rest of this post walks through what we tested, what changed my mind, and how I would run a similar test inside another engineering organization.
Can you trust public AI benchmarks for choosing a coding model?
The public signal is strong enough that software leaders should pay attention. Open models are closing benchmark gaps quickly, and recently released open models such as GLM-5.2, Kimi K2.7, Qwen3-Coder, and DeepSeek V4 are explicitly aimed at long-context, tool-using, agentic coding workflows.
Those benchmarks set priors. They cannot see your repositories, test setup, review standards, cache behavior, CI failure modes, permission boundaries, or the mix of work your engineers actually need help with. A model can look strong on a leaderboard and still be the wrong default for your environment if the harness wastes context, the cache economics break, or the model struggles on your common task shapes.
I wanted a Faros-specific answer. Which model and harness should we trust for our repositories, our task mix, and our review standards?
This is related to the new wave of model routers, including Sakana's Fugu, which dynamically orchestrates multiple models behind a single API. I like that direction. The limit is that a generic router still needs local evidence: the judge and routing policy have to reflect your codebase, tests, permission boundaries, review burden, cache behavior, and the work slices your engineers actually do.
How we tested different AI coding models on real engineering tasks
For our experiment, we compared model plus harness combinations. I care about the pairing because engineers do not use a raw model in isolation. They use a model through a harness that decides how context is loaded, how tools are called, how diffs are produced, and how failures show up. I wanted the methodology visible before the result, because this experiment is only useful if the reader knows what kind of evidence it is.
Question
Experiment detail
What was the main question the experiment intended to answer?
Could an open-model route handle real Faros engineering work well enough that the frontier route should stop being the default for some tasks?
What work source did we use?
Historical Faros engineering tasks and PR context, transformed into Time Machine task rows. (This is local company evidence, not a public benchmark.)
How were tasks selected?
The source pool was 500 selected historical Faros PR/Jira instances across 12 repositories. The headline comparison uses the 211 rows that survived the strict cohort filter.
Were they sampled randomly, stratified, or handpicked?
The cohort was curated for diversity and operational representativeness. We selected 500 task<>PR pairs from the most active Faros repositories, with coverage across complexity levels and work types: bugs, feature work, KTLO/maintenance, infra/devex, migrations, code-review-style work, and several product surfaces. The selector then balanced within each repo across category and size buckets and filtered out unusable rows, such as missing base commits, missing gold patches, etc.
What was the main cohort?
211 strict instances across 12 Faros repositories. Work type: 61 KTLO/maintenance tasks, 54 feature tasks, 49 infra/devex tasks, 27 bug fixes, and 20 other tasks. Complexity: 81 very-high-complexity tasks, 61 high-complexity tasks, 50 medium-complexity tasks, and 19 low-complexity tasks. Repository domain: 87 UI/reporting tasks, 67 data graph/platform tasks, 47 AI/agent/ML tasks, and 10 connector-ingestion tasks.
What routes did we compare?
OK: OpenCode + Kimi K2.6 — Fireworks default CK: Claude Code + Kimi K2.6 — Fireworks default OG: OpenCode + GLM-5.2 — Fireworks default, Max reasoning CG: Claude Code + GLM-5.2 — Fireworks default, Max reasoning OO: OpenCode + Opus 4.8 — default high effort CO: Claude Code + Opus 4.8 — default high effort Codex: Codex + GPT-5.5 — configured high reasoning
Did each harness receive equivalent context and tool permissions?
Each route received the same task setup: the same historical repo state, task/PR context, prompt shape, and permission to produce a patch. What varied was the route itself: the harness/model pairing. So this is not a raw model benchmark; it measures the practical route an engineer would actually use.
Were routes run blind?
The implementation routes were blind to the historical solution. Each route received the same scrubbed task context and repo state at the base commit; it did not receive the accepted PR diff, test patch, PR number, merge metadata, author, issue hints, or evaluator-only fields. No route received extra manual help or additional context beyond the shared setup.
Did the judge see model identity?
The scoring code used a model-based rubric judge over the task description, generated patch, and per-task rubric. The judge prompt did not include the route, harness, or implementation model identity; those labels were retained separately as experiment metadata/provenance.
Was the judge model-based, human-reviewed, test-based, or composite?
The headline judge score was model-based rubric scoring.
How were invalid/no-diff cases handled outside the strict cohort?
They were tracked separately in implementation telemetry as invalid, no-diff, failure, or timeout cases. They were not mixed into the headline mean-score table.
What was excluded by the strict cohort filter?
Rows were excluded from the headline cohort if any selected seven-way route lacked a valid implementation or judge score.
Were costs actual billed costs, projected normalized costs, or replay-estimated costs?
Projected normalized/list costs from token telemetry and provider pricing.
Were retries allowed?
Gap-fix and recovery reruns were used where needed to repair telemetry or completion gaps. During the experiment, we found a caching issue that was specific to the Claude Code + Kimi route. We worked with the Fireworks team to resolve it, then reran that route and used the cache-fixed rerun in the final comparison, excluding the earlier low-cache CK run.
Were routes run once or multiple times?
The headline table uses one implementation and one score per route/task, not an average over repeated trials. Some selected rows came from rerun or gap-fix sources; the run source is retained in the appendix.
What outcomes did we measure?
Aggregate judge score, task win rate, average rank score, projected cost per task, runtime, cache share, generated patch bytes, and invalid/no-diff behavior.
What definitions matter for reading the results?
Time Machine: Faros analysis for replaying historical engineering work. Harness: coding-agent wrapper around the model. Strict cohort: every selected route had a valid implementation and score. Judge score: model-based rubric score. Projected cost: token telemetry plus provider pricing. Provider default: Fireworks-hosted Kimi and GLM rows used provider/model defaults. Cache share: share of input context served from cache.
What should the reader keep in mind?
This is a production-style routing evaluation. It is useful for deciding which route to test as a default, but it should not be read as a universal model leaderboard.
Questions and details for interpreting Faros's routing evaluation experiment.
Cohort shape: the task, complexity, and repository-domain mix behind the result.
GLM-5.2 vs. Kimi K2.6 vs. Claude Opus: Quality and cost results
GLM-5.2 and Kimi both landed in the top quality band. Claude Code + GLM-5.2 scored 0.568 and Claude Code + Kimi scored 0.566, close enough that I would not treat the difference as a quality win by itself. The operational difference was economics: in this cohort, GLM-5.2 delivered comparable quality while running faster and costing meaningfully less per task. That makes GLM-5.2 the route I would test first for broad production-like throughput, with Kimi still worth keeping in the routing pool where local data shows an edge.
Variant
Judge score
Cache share
Avg time/task
Avg cost/task
CG
0.568
89.7%
321s
$0.92
CK
0.566
90.0%
386s
$1.78
OG
0.556
95.5%
620s
$1.01
CO
0.521
99.7%
775s
$1.76
OK
0.514
94.8%
557s
$0.86
Codex
0.466
92.8%
392s
$2.06
OO
0.454
100.0%
544s
$1.05
Findings from model routing comparison experiment
Quality vs. projected cost: the seven-way Pareto view shows why Claude Code + GLM-5.2 became the route I would test first for broad throughput.
Because cache share was nearly identical after the rerun, 89.7% for Claude Code + GLM-5.2 and 90.0% for Claude Code + Kimi, the cost gap was not a cache artifact.
Quality vs. runtime: the Pareto view shows Claude Code + GLM-5.2 as the quality-and-speed frontier, with Claude Code + Kimi close on quality but slower on average.
The expensive frontier baselines did not buy their way into the top quality band in this cohort. Claude Code + Opus 4.8 trailed both Claude Code + Kimi and Claude Code + GLM-5.2 on mean score; its projected cost was roughly in the same range as Claude Code + Kimi, but materially higher than Claude Code + GLM-5.2. Codex + GPT-5.5 is included here as a high-reasoning current-default comparison, not as a Fireworks route.
How should you route AI coding tasks by work type?
The aggregate table tells me which routes are efficient overall. The work-type slices tell me how I would route real work. I would not turn the aggregate winner into a one-size-fits-all rule.
Mean score tells me average quality. Average rank tells me which route stayed consistently near the top across tasks, even when it did not win outright.
Consistency vs. projected cost: the Pareto view shows OpenCode + GLM-5.2's average-rank strength while keeping Claude Code + GLM-5.2 in view as the broad default candidate.
For this cohort, the routing policy would start here:
Work type
Start here
Escalate when
Signal from this cohort
Bugs (27)
Claude Code + GLM-5.2
Claude Code + Kimi for subtle or high-blast-radius fixes
CG had the strongest bug readout: top mean score, top win rate, and top rank.
Feature work (54)
Claude Code + GLM-5.2 for routine throughput
Claude Code + Kimi for ambiguity, larger diffs, or review risk
CK led this slice by mean and rank.
Infra/devex (49)
Claude Code + Kimi when correctness dominates; OpenCode + GLM-5.2 when consistency matters
OpenCode + Kimi for lower-risk work where cost discipline matters most
Split signal: CK led by mean; OG led by rank.
KTLO / maintenance (61)
Claude Code + GLM-5.2
Claude Code + Kimi when blast radius, dependencies, or review churn are high
CG led by mean.
Data graph / platform (67; 40 schema-graph dominant)
OpenCode + GLM-5.2
Claude Code + Kimi or Claude Code + GLM-5.2 when review risk dominates
OG was strongest, especially on schema-heavy context.
Claude Code + GLM-5.2 only after local validation on similar tasks
This slice is closer to the hardest work; subtle failures matter more than cost.
Example routing policy
The point is not to memorize these labels. The operating rule is to start with Claude Code + GLM-5.2 where local data says it clears the bar, escalate to Claude Code + Kimi when ambiguity, larger diffs, complexity, or review risk make quality worth paying for, and keep OpenCode + GLM-5.2 as the OpenCode route with the strongest consistency signal.
GLM-5.2 for coding: Why it changed our default model
GLM-5.2 was announced while we were running this experiment, so we added it to the loop. The GLM-5.2 release post emphasized long-horizon engineering, larger context, and stronger coding benchmark performance. That made it worth testing. The completed run made it practical: Claude Code + GLM-5.2 landed in the top quality band, ran faster than Claude Code + Kimi, and cost roughly half as much per task.
This is what I want to emphasize: The best default changed while the experiment was still running. Kimi still remains the route I would keep for quality-sensitive escalation, especially where the local slices show an edge: feature work, very-high-complexity work, larger diffs, and agent-tooling tasks. But GLM-5.2 cleared the bar often enough—and cheaply enough—that it changed what I would test first.
OpenCode + GLM-5.2 also earned its place. It had the best average rank signal, strong cache behavior, and the best data graph / schema-heavy profile. If a team is standardizing around OpenCode, that route deserves serious evaluation rather than being treated as a fallback.
The point is not to pick a winner and freeze it. You need the ability to rerun the evaluation when the market moves, when a provider fixes caching, when a new model ships, or when your own workload changes. This has to become an operating loop and not a one-time bakeoff.
How to run an AI model routing experiment (step-by-step)
Model routing is becoming an important engineering control. If I were doing this experiment with another engineering team, I would start small and make the decision useful within a week. Running your own small experiment could look like this:
Step
What to measure
Decision it informs
Pick a representative sample
20-50 recent tasks across bugs, features, infra/devex, KTLO, and maintenance
Keeps the result grounded in your actual work mix instead of treating a public leaderboard as a proxy for your codebase
Run two or three routes
A frontier route, an open-model route, and one alternate harness if available
Shows whether the default should change
Score quality and review burden
Patch quality, test impact, reviewer effort, and failure modes
Separates cheap from actually useful
Track economics
Projected cost, cache share, runtime, retries, and invalid/no-diff outputs
Finds where savings survive real execution
Turn it into a policy
Task types where the cheaper route clears the bar, escalation rules, and rerun cadence
Makes model choice repeatable instead of anecdotal
Evaluation steps, measurements, and decisions for comparing model routes.
Focus your model routing experiment on real tasks like test selection, deploy gates, and review policies. The process should reflect your actual workflows, adapt to changing conditions, and define exactly when to escalate. The goal is not to crown a winning model. The goal is to make model choice repeatable.
{{cta}}
Why AI model selection should be a continuous loop
Defaults need evidence. Open models now deserve real engineering evaluations on actual company work, turning model choice into an engineering control like test selection or deploy gates. The market simply moves too quickly for static defaults. When a new model ships, caching improves, or your workload shifts, the right answer changes. Therefore, your most durable capability is continuously measuring your own work to adapt. Model selection must be an operating loop, not a one-time bake-off.
Faros can produce model-routing policies from the engineering systems you already use, drawing on issues, PRs, CI, code ownership, review flow, repo metadata, and delivery outcomes. Time Machine turns that history into a clear guide for which route should handle which work, when to escalate, and where the economics actually survive contact with your codebase. Contact us for a demo today.
Frequently asked questions when comparing open source vs. frontier AI models for coding
Are open source AI models as good as GPT-5.5 or Claude Opus for coding?
Yes, for certain engineering work. In Faros's test of 211 real coding tasks, open models GLM-5.2 and Kimi K2.6 scored in the top quality band, while frontier routes using Claude Opus 4.8 and GPT-5.5 scored lower and cost more. GLM-5.2 matched top quality at roughly half the cost per task.
What is AI model routing?
AI model routing is the practice of directing each engineering task to the model best suited for it, instead of sending every request to one expensive frontier model. Faros found this can cut AI coding costs by roughly 50% at equal quality, because cheaper open models handle routine work while costlier models are reserved for high-risk tasks.
How much can teams save by switching from frontier models to open models for coding?
Roughly 50% per task. In Faros's test, the best open-model route (GLM-5.2) cost $0.92 per task versus $1.76–$2.06 for frontier routes using Claude Opus 4.8 and GPT-5.5—at comparable quality. Actual savings depend on task mix and how much work can safely use the cheaper route.
What are the best open source AI models for coding in 2026?
The leading open source coding models in 2026 include GLM-5.2, Kimi K2.6, Qwen3-Coder, and DeepSeek V4, all built for long-context, agentic workflows. In Faros's testing of 211 tasks, GLM-5.2 and Kimi K2.6 performed strongest on real engineering work, with GLM-5.2 offering the best quality-to-cost ratio.
Can you trust public AI benchmarks when choosing a coding model?
Not fully. Public benchmarks are useful priors but can't account for your codebase, review standards, cache economics, or CI failure modes. Faros's test showed a top-benchmark model can still be the wrong default for a given environment. Validate candidate models against a sample of your own tasks first.
Ron Meldiner
Ron is an experienced engineering leader and developer productivity specialist. Prior to his current role as Field CTO at Faros, Ron led developer infrastructure at Dropbox.
Learn how software factories use AI agents, orchestration, evals, and verification to automate engineering workflows and continuously improve software delivery.
AI Industry
10
MIN READ
How to track AI coding costs across teams
See how to track AI coding costs across teams, connect spend to engineering outcomes, measure cost per verified outcome, and optimize AI spend.
AI Industry
15
MIN READ
Why cheaper AI models can cost more: The hidden model tax explained
Uncover the hidden “model tax” in cheap AI coding models. Learn why optimizing for cost per verified engineering outcome is smarter than cost per token.