AI Economics

What Did One Successful AI Outcome Actually Cost?

A framework for measuring cost per successful AI outcome across attribution boundaries, retries, delayed success evidence, and changing economic definitions.

Licenzy TeamPublished Oct 6, 202622 min read

What Did One Successful AI Outcome Actually Cost?

Consider two AI workflows.

Both process the same number of executions. Both accumulate the same attributable cost.

MeasureWorkflow AWorkflow B
Executions100100
Attributable cost$100$100
Cost per execution$1.00$1.00

From an execution-cost perspective, they look identical.

Each workflow processed 100 executions at an average attributable cost of $1.00 per execution.

Now add one more piece of information.

MeasureWorkflow AWorkflow B
Executions100100
Attributable cost$100$100
Successful outcomes9050
Cost per execution$1.00$1.00
Cost per successful outcome$1.11$2.00

Nothing changed about the execution count. Nothing changed about the attributable cost in this example. Nothing changed about the cost per execution.

But the two workflows did not produce the same number of outcomes classified as successful.

Relative to those successful outcomes, the same $100 of attributable cost now tells a different story.

Same spend. Same execution count. Same cost per execution. Different outcome economics.

That does not mean Workflow B is automatically worse. It does not tell us which workflow is more profitable. It does not tell us which outcomes created more customer value.

And it does not tell us whether either workflow should be priced differently.

We do not have enough information to answer those questions.

What the comparison does show is narrower:

Stable cost per execution does not necessarily mean stable cost to produce successful outcomes.

That gives us a different economic question to investigate.

What did one successful AI outcome actually cost?

Cost per Execution and Cost per Successful Outcome Answer Different Questions

Cost per execution is useful.

If an AI workflow consumes $100 of attributable cost across 100 executions, an average of $1.00 per execution tells us something meaningful about the cost of its runtime activity.

That can help us reason about execution behavior, compare versions of a workflow, investigate infrastructure or model changes, and understand how expensive it is to run the system.

The problem begins only when we ask that metric to answer a different question.

Suppose two workflows both cost:

text
$1.00 per execution

but one produces a successful outcome in 90 of 100 executions while the other produces one in 50.

The execution-level metric remains the same. The relationship between economic work and successful results does not.

QuestionMeaning
EXECUTION QUESTIONHow expensive is the activity?
OUTCOME QUESTIONHow much attributable cost corresponds to the outcomes classified as successful?

One does not replace the other.

The first keeps the denominator close to runtime activity. The second relates cost to a result we have decided counts as successful.

That distinction matters especially in AI systems because a single result may involve variable model usage, retrieval, tool calls, external APIs, validation, retries, or other runtime work—and because not every execution necessarily produces the outcome the product cares about.

So a workflow can become cheaper to execute without necessarily becoming cheaper at producing successful results.

The reverse can also happen. Execution cost can remain stable while the number of successful outcomes changes.

Execution cost measures activity. Outcome cost relates economic work to successful results.

But that second view creates two problems that the arithmetic alone cannot solve.

What exactly belongs in the cost?

And what exactly qualifies as a successful outcome?

Those questions sit on opposite sides of a deceptively simple formula.

The Formula Is Simple. The Numerator and Denominator Are Not.

At its simplest, the relationship looks like this:

text
             ATTRIBUTABLE COST
        ─────────────────────────
           SUCCESSFUL OUTCOMES

The division is easy.

The definitions underneath it are not.

The numerator requires us to decide which economic work belongs inside the analysis. The denominator requires us to decide which outcomes qualify as successful.

If either side changes meaning, the resulting number changes meaning with it.

We will use cost per successful outcome descriptively for this relationship.

Related ideas already appear in AI systems and benchmarks as cost per successful task, cost per resolution, and cost per outcome. But those terms do not imply one universal definition of either cost or success.

That distinction matters.

A formula can produce a number with two decimal places while still hiding disagreement about what was measured.

The formula is simple. The numerator and denominator are not.

So before interpreting the result, we need to open both sides of the fraction.

Start with the numerator.

The Numerator Is an Attribution Decision

Imagine an AI coding workflow producing Outcome #123.

One outcome can involve more than one economic operation:

Relationship figureOutcome #123 runtime flow
  • Initial execution
  • Model inference
  • Retrieval
  • Tool call
  • Validation
    • FAILED
  • Retry
  • Additional model call
  • Validation
    • PASSED
  • Successful outcome

Now ask a seemingly simple question:

What did Outcome #123 cost?

Model, retrieval, tools, validation, failed paths, and shared infrastructure may all be relevant. Simply dividing every company-level AI expense by successful outcomes would not automatically produce a meaningful numerator.

For this Guide, attributable cost means the cost included within the analytical boundary chosen for the population of outcomes being measured. That boundary can include directly observed execution cost, defensible allocations of shared cost, and estimates; other costs may remain unknown or intentionally excluded. Those are not equivalent forms of evidence, and the metric should not pretend otherwise.

Consider two calculations:

text
Analysis A

Model inference
+ retrieval
+ external tools

and:

text
Analysis B

Model inference
+ retrieval
+ external tools
+ allocated infrastructure
+ human review

Both can be valid for a clearly defined purpose, but they are not automatically comparable merely because both are labeled:

text
cost per successful outcome

Their numerators describe different cost boundaries. Total AI spend is not automatically the right numerator merely because it is easiest to obtain; the population above the line has to make analytical sense relative to the population below it.

And even after that boundary is defined, another complication remains.

Some of the economic work associated with a successful outcome may not itself have succeeded.

The Successful Attempt Is Not Necessarily the Whole Economic Story

Suppose Outcome #123 required two attempts.

text
Attempt 1
FAILED
$0.16

Attempt 2
SUCCESS
$0.18

The successful attempt cost $0.18. But that does not automatically answer a different question:

What did it economically take to produce the successful outcome?

If the second attempt exists because the first one failed, it may be tempting to write:

text
$0.16 + $0.18
=
$0.34 per successful outcome

That arithmetic is correct.

The attribution is not automatically established.

Depending on the analytical question, the first attempt may be attributable to the eventual outcome, treated as reliability overhead, allocated differently, or have an uncertain relationship to it.

So we can say:

text
Successful attempt cost
=
$0.18

But before saying:

text
Cost of producing the successful outcome
=
$0.34

we need an attribution decision.

The cost of the successful attempt is not necessarily the cost of producing the successful outcome.

The important point is not that every failed attempt belongs to the successful outcome that followed it. It is that a failed path can remain economically relevant while adding nothing to the successful-outcome count.

Failed work can affect the numerator without increasing the denominator.

That is one reason two systems with the same number of successful outcomes can require different amounts of economic work to produce them.

But solving the numerator is only half of the problem.

Even if we could perfectly reconstruct every relevant unit of cost, we would still need to decide which outcomes are allowed below the line.

The Denominator Is a Success Decision

Suppose an AI system records:

MeasureCount
Executions1,000
Completed executions850
Outputs passing validation700
Outcomes satisfying the chosen success condition600

Which number belongs in the denominator?

There is no universal answer.

We could calculate:

  • cost per execution
  • cost per completed execution
  • cost per validated output
  • cost per successful outcome

Each ratio answers a different question.

If we want to understand the economics of execution activity, executions may be the relevant denominator. If we want to understand the economics of outputs that passed a specific validation rule, validated outputs may be appropriate.

If we want to understand the economic work associated with outcomes that satisfied a defined success condition, then successful outcomes become relevant.

The last denominator is not automatically superior. It is simply aligned with a different question.

The denominator should match the economic question.

But once we choose successful outcomes, we inherit another requirement.

We need to know what successful means.

For one coding agent, success might mean:

text
required tests passed

For another analysis, it might mean:

text
change accepted into the codebase

For another:

text
reported production problem resolved

Those conditions are not interchangeable.

And the evidence required to support them may not arrive from the same place or at the same time.

A deterministic validation can establish one kind of condition. A domain event can establish another. Human review, user behavior, or a model-based evaluation may provide evidence for others.

The purpose here is not to choose one universal evidence hierarchy.

It is to make the denominator explicit.

The denominator is only as defensible as the rule and evidence used to classify outcomes as successful.

A number called successful outcomes can look precise while still hiding a changing definition underneath it.

And there is another complication.

Sometimes we do know what success means.

We simply do not know the answer yet.

Pending Is Not Failed

Suppose 100 AI executions have completed.

At the moment we inspect them, their outcome state looks like this:

MeasureCount
Executions100
SUCCESSFUL70
FAILED10
PENDING20

Assume the attributable cost for those executions is $100.

If we divide that cost by the 70 outcomes currently confirmed as successful, we get:

text
$100 / 70
≈ $1.43 per confirmed successful outcome

That calculation may be correct for the information currently available.

But it does not mean the other 30 outcomes all failed.

Ten are classified as failed. Twenty are still pending.

That distinction matters when success depends on evidence that arrives after execution.

A coding agent may complete before a review is approved. A support workflow may respond before the customer confirms whether the issue was resolved. An AI-generated recommendation may be produced before the downstream application event that determines whether the intended outcome occurred.

Now imagine that seven days later the same population looks like this:

StateDAY 1DAY 7
SUCCESSFUL7085
FAILED1010
PENDING205

No new execution needs to have occurred for the denominator to change.

What changed was the available outcome evidence.

Using the same illustrative $100 cost boundary:

text
DAY 1

$100 / 70
≈ $1.43 per confirmed successful outcome

DAY 7

$100 / 85
≈ $1.18 per confirmed successful outcome

The apparent economics changed even though the execution population and the cost assigned to it did not.

The denominator can change because evidence arrived, not because new execution occurred.

This is why PENDING should not automatically be collapsed into FAILED.

And it is why the observation window matters.

If the success condition depends on delayed evidence, a recent cohort may contain outcomes whose final classification is not yet known.

That does not make the metric unusable. It means the metric needs context.

A cost-per-success figure based only on currently confirmed outcomes describes what is known at that point in time.

It should not silently pretend that unresolved outcomes have already reached their final state.

Cheap Execution Does Not Necessarily Mean Cheap Successful Work

We can now return to the two workflows from the beginning.

MeasureWorkflow AWorkflow B
Executions100100
Attributable cost$100$100
Successful outcomes9050
Cost per execution$1.00$1.00
Cost per successful outcome$1.11$2.00

Their execution-cost profiles are identical in this simplified example.

Their observed outcome yield is not.

Workflow A produced 90 outcomes classified as successful from the same 100 executions and $100 of attributable cost. Workflow B produced 50.

That difference becomes invisible if we look only at cost per execution.

This does not make cost per execution wrong. It means cost per execution is answering the question it was built to answer.

It tells us about the economics of execution activity. It does not incorporate how often that activity produces the result we have chosen to count as successful.

Cheap execution and cheap successful work are different properties.

This distinction is already appearing in contemporary AI-agent benchmarking.

A 2026 benchmark published by Arize AI and Fireworks evaluated 2,400 agent runs across 10 models and 40 tasks using, among other measures, cost per successful task. Its methodology calculates that measure from spend across attempts relative to successful runs, making success rate and attempt cost jointly relevant to the comparison.

That can change how systems compare.

A model or agent that is inexpensive per attempt can still require more economic work per successful task if it succeeds less often or requires additional attempts.

Conversely, a more expensive execution can sometimes correspond to a lower cost per successful task if it produces successful results more reliably.

The important lesson is not that one metric should replace the other.

It is that the comparison can depend on the question.

QuestionMeaning
EXECUTION QUESTIONHow expensive is the activity?
OUTCOME QUESTIONHow much economic work corresponds to the successful results?

Those questions become especially important in agentic systems because the path to one outcome may vary substantially from another.

One execution may complete in a single model call. Another may require retrieval, multiple tools, validation, retries, or additional model calls.

Two systems can therefore look similar at one analytical boundary and different at another.

But even if we build a defensible measure of cost per successful outcome, we still have not answered the most important business question.

We know something about the cost of producing success.

We do not yet know what that success was worth.

Cost per Successful Outcome Is Not Value or Profitability

Suppose two AI workflows each produce 100 successful outcomes with $100 of attributable cost.

MeasureWorkflow AWorkflow B
Successful outcomes100100
Attributable cost$100$100
Cost per successful outcome$1.00$1.00

At this analytical boundary, their cost per successful outcome is identical.

But that does not mean the outcomes are economically equivalent.

One workflow might produce a relatively low-value result. Another might complete work that customers value much more highly. Their commercial models may differ.

Their revenue may differ. And other relevant costs may sit outside the boundary used in this calculation.

The ratio does not tell us any of that.

text
COST PER SUCCESSFUL OUTCOME
            ≠
      VALUE PER OUTCOME

A successful outcome is not automatically a unit of customer value.

It is an outcome that satisfied the success condition chosen for the analysis.

That condition might be highly correlated with customer value. Or it might represent only one intermediate step toward it.

For example, an AI coding agent might successfully produce a change that passes the required tests.

That is a meaningful success condition.

But it does not, by itself, tell us how valuable the change is to the customer, how much revenue it supports, or whether it ultimately improves the business outcome the customer cares about.

This distinction also prevents another tempting shortcut:

text
LOW COST PER SUCCESSFUL OUTCOME
                ≠
            PROFITABLE

Imagine:

MeasureWorkflow AWorkflow B
Cost per successful outcome$0.50$4.00
Revenue associated with it$0.30$20.00

These numbers are deliberately simplified, and they are not a complete profitability calculation.

But they expose the problem.

Workflow A has the lower cost per successful outcome. That alone does not make it economically better.

Workflow B has a much higher cost per successful outcome, yet the commercial context around that outcome is also very different.

Profitability requires more information than the cost side of the ratio can provide.

Depending on the question, that may include revenue, commercial terms, additional cost categories, customer or workload mix, and the economic boundary being analyzed.

So there are at least three different questions hiding near each other:

QuestionMeaning
RUNTIME QUESTIONWhat does execution activity cost?
OUTCOME QUESTIONWhat does it cost to produce results classified as successful?
BUSINESS QUESTIONWhat are those successful outcomes worth in their commercial and economic context?

These are not a universal hierarchy.

They are different analytical questions.

A system can improve on one while remaining unchanged—or even moving in the opposite direction—on another.

Cost per successful outcome tells us something about the cost side of producing success. It does not tell us what that success was worth.

And keeping those questions separate matters for another reason.

Measuring economics around an outcome does not mean the outcome has to become the thing the customer buys.

Measuring per Outcome Does Not Mean Charging per Outcome

Suppose an AI product internally measures:

text
$1.20 attributable cost
per successful outcome

Nothing about that measurement determines how the product must be sold.

The commercial model could still be:

text
Subscription

Credits

Usage-based pricing

Seats

Outcome-based pricing

or a combination of these

The commercial model can still be a subscription, credits, usage-based pricing, seats, outcome-based pricing, or a combination. Real products already use concepts such as resolutions or outcomes in operational reporting and, in some cases, commercial measurement.

For example, Intercom defines multiple Fin AI Agent outcome types and uses resolutions within its commercial model. Zendesk similarly uses automated resolutions as a measure of AI-agent usage.

But the existence of outcome-based commercial models does not turn cost per successful outcome into a pricing prescription.

These are separate decisions.

Relationship figureInternal economic view and commercial model
  • INTERNAL ECONOMIC VIEW
    • How much attributable cost corresponds to successful outcomes?
  • ≠
  • COMMERCIAL MODEL
    • What does the customer buy, consume, or get charged for?

A company can sell a subscription, deduct credits, charge for usage, or use an outcome-based commercial unit while separately analyzing the cost of each successful outcome.

Measuring economics per successful outcome does not require selling the product per successful outcome.

Customers do not need to absorb the complexity of the runtime for the business to measure it. Conversely, an outcome-based price does not automatically tell the business what that outcome cost to produce.

Suppose a customer is charged:

text
$5 per resolved task

That commercial fact does not reveal whether the attributable cost behind one resolution was:

text
$0.40

or:

text
$4.50

or whether the cost varies substantially across different kinds of tasks.

Commercial measurement and economic measurement can intersect without becoming the same thing.

That brings us to a final danger.

Once a metric feels closer to the outcome the product cares about, it is easy to treat it as the number that finally tells us whether the system is improving.

It does not.

A better denominator can still produce a metric we misunderstand.

Don't Turn It Into Another Vanity Metric

Suppose your dashboard shows this:

text
Cost per successful outcome

August       $1.80
September    $1.20

Change       -33%

At first glance, that looks like an improvement.

Perhaps it is.

But the metric itself does not tell us why it changed.

The underlying execution may have become cheaper. The success rate may have improved. The workflow may require fewer retries. A provider rate may have changed.

The mix of workloads may be different. The attribution boundary may have changed. The success condition may have changed.

Or the observation window may contain a different proportion of outcomes whose success is still pending.

Several of these changes can happen at the same time.

So even after moving from execution-level cost toward outcome-level economics, an old analytical problem remains:

A change in cost per successful outcome is an observation before it is an explanation.

Consider two very different scenarios.

ScenarioAttributable costSuccessful outcomesCost per successful outcome
Scenario A$180 → $120100 → 100$1.80 → $1.20
Scenario B$180 → $180100 → 150$1.80 → $1.20

Here, the measured improvement comes from the numerator.

Now consider Scenario B.

The resulting metric is identical.

The mechanism underneath it is not.

In Scenario A, less attributable cost corresponds to the same number of successful outcomes. In Scenario B, the same attributable cost corresponds to more successful outcomes.

Those are economically different stories hidden behind the same final ratio.

And there are more possibilities.

A lower provider rate could reduce the numerator without changing runtime behavior. A model change could increase the number of successful outcomes while also increasing cost per execution.

A retry-policy change could alter both cost and success yield. A shift toward easier workloads could improve the aggregate metric even if the economics of each workload remained unchanged.

Or nothing about the underlying system may have improved at all.

The measurement itself may have changed.

Comparisons Require Stable Meaning

Suppose that in August an outcome counted as successful when:

text
validation passed

but in September the rule became:

text
user accepted the result

Now compare:

text
August       $1.80 per successful outcome
September    $1.20 per successful outcome

The arithmetic still works.

The comparison may not mean what it appears to mean.

The denominator no longer represents the same success condition.

The same problem can occur above the line.

Perhaps August includes only directly observed model and tool costs, while September also includes allocated infrastructure. Or perhaps the valuation basis changed because the applicable provider rates changed.

Or one period gives outcomes seven days to resolve while another measures them after only twenty-four hours.

When comparing cost per successful outcome across periods, workflows, models, or customers, interpretation therefore depends on knowing whether important parts of the measurement remained comparable.

That includes, among other things:

text
cost boundary

success definition

valuation basis

observation window

These do not have to remain frozen forever.

Definitions may change when the analytical question changes or the measurement improves.

But those changes need to be visible.

Otherwise, a measurement change can masquerade as an economic change.

A metric can change because the system changed—or because the meaning of the metric changed.

Even a consistently defined aggregate can hide another problem.

LevelCost per successful outcome
Aggregate$1.40
Workflow A$0.80
Workflow B$1.20
Workflow C$4.70

The aggregate can be useful.

It is also incomplete.

A change in workload mix could move the aggregate even if none of the individual workflows became more or less efficient at producing successful outcomes.

Aggregate outcome economics can hide materially different workflow economics.

That does not mean every metric needs to be segmented endlessly.

It means the level of aggregation should match the question being investigated.

The same principle has followed us throughout this Guide.

The number is not the explanation.

A More Outcome-Relevant Denominator Does Not Remove the Need for Explanation

Cost per execution can hide differences in success yield.

Cost per successful outcome can expose some of those differences.

That makes it useful.

It does not make it self-explanatory.

Imagine the metric rises next month:

text
$1.20 → $1.65

What happened?

Perhaps executions became more expensive. Perhaps executions cost exactly the same, but fewer produced successful outcomes. Perhaps retries increased.

Perhaps the underlying resource rates changed. Perhaps the workload mix changed. Perhaps more outcomes remained pending at measurement time.

Perhaps the attribution boundary changed. Perhaps the definition of success became stricter.

Looking only at the final ratio cannot distinguish among those explanations.

To understand the change, we still need evidence from underneath the metric.

Execution evidence.

Cost evidence.

Attribution evidence.

Outcome evidence.

And commercial context when the question extends beyond cost.

A more outcome-relevant denominator does not remove the need for explanation.

This matters because improving the metric without understanding the mechanism can be just as misleading as optimizing the wrong metric.

A lower cost per successful outcome may be desirable.

But before treating it as evidence that the system became economically better, we need to know what produced the change.

The metric tells us that the measured relationship changed.

Attribution can help us locate where the change occurred.

The underlying evidence helps us explain what changed.

What Did Success Actually Cost?

We started with two workflows that looked identical from one perspective.

MeasureWorkflow AWorkflow B
Executions100100
Attributable cost$100$100
Cost per execution$1.00$1.00

Then we added the outcome:

MeasureWorkflow AWorkflow B
Successful outcomes9050
Cost per successful outcome$1.11$2.00

The arithmetic revealed something that cost per execution could not.

But the arithmetic was never the difficult part.

To interpret the numerator, we needed to know what economic work belonged inside the boundary. To interpret the denominator, we needed to know what counted as success and what evidence supported that classification.

Failed attempts could consume resources without becoming successful outcomes. Retries could increase the economic work associated with reaching one. Pending outcomes could leave the denominator temporarily incomplete.

And even after all of that, cost per successful outcome still could not tell us what the outcome was worth, whether the product was profitable, or how the customer should be charged.

Those are different questions.

The useful distinction is therefore not:

LabelMetric
BAD METRICcost per execution
GOOD METRICcost per successful outcome

It is:

Relationship figureRuntime activity and successful results
  • RUNTIME ACTIVITY
    • What did execution cost?
  • different question
  • SUCCESSFUL RESULTS
    • What economic work did it take to produce the outcomes we count as successful?

Both views can matter.

They simply illuminate different boundaries of the system.

For AI products, that distinction becomes particularly useful because the path between execution and outcome can contain variable model usage, retrieval, tools, external services, validation, retries, and uncertain or delayed success evidence.

The same commercial product can therefore produce the same number of executions while requiring very different amounts of economic work to produce successful results.

And the same cost per successful outcome can emerge from very different mechanisms underneath.

So the goal is not to find one number that finally summarizes AI economics.

It is to make the economic question explicit enough that the number has a defensible meaning.

The formula is simple. The numerator and denominator are not.

And once both sides are defensible, one question remains:

If successful outcomes became more expensive, do you know whether the cost changed, the success rate changed—or your definition did?

References

Arize AI & Fireworks AI — Cost per Successful Task Benchmark — Benchmark of AI-agent economics across 2,400 runs, 10 models, and 40 tasks, including cost per successful task and the economic effect of attempts that do not successfully complete the benchmark task.

Arize AI — Fireworks Cost Benchmark Repository — Public benchmark methodology and implementation, including the calculation of cost per successful task from total spend across attempts and successful runs.

FinOps Foundation — Allocation — Framework for allocating technology costs, including directly attributable and shared costs and the use of allocation methodologies where costs cannot be mapped directly to a single unit.

Intercom — Fin AI Agent Outcomes — Documentation describing Fin outcome types, resolution rules, and the operational and commercial treatment of AI-agent outcomes.

Zendesk — About Automated Resolution Tiers — Documentation describing automated resolutions as a measure of AI-agent usage and their role in Zendesk's commercial model.

Related Guides