What Did One Successful AI Outcome Actually Cost?
A framework for measuring cost per successful AI outcome across attribution boundaries, retries, delayed success evidence, and changing economic definitions.
What Did One Successful AI Outcome Actually Cost?
Consider two AI workflows.
Both process the same number of executions. Both accumulate the same attributable cost.
| Measure | Workflow A | Workflow B |
|---|---|---|
| Executions | 100 | 100 |
| Attributable cost | $100 | $100 |
| Cost per execution | $1.00 | $1.00 |
From an execution-cost perspective, they look identical.
Each workflow processed 100 executions at an average attributable cost of $1.00 per execution.
Now add one more piece of information.
| Measure | Workflow A | Workflow B |
|---|---|---|
| Executions | 100 | 100 |
| Attributable cost | $100 | $100 |
| Successful outcomes | 90 | 50 |
| Cost per execution | $1.00 | $1.00 |
| Cost per successful outcome | $1.11 | $2.00 |
Nothing changed about the execution count. Nothing changed about the attributable cost in this example. Nothing changed about the cost per execution.
But the two workflows did not produce the same number of outcomes classified as successful.
Relative to those successful outcomes, the same $100 of attributable cost now tells a different story.
Same spend. Same execution count. Same cost per execution. Different outcome economics.
That does not mean Workflow B is automatically worse. It does not tell us which workflow is more profitable. It does not tell us which outcomes created more customer value.
And it does not tell us whether either workflow should be priced differently.
We do not have enough information to answer those questions.
What the comparison does show is narrower:
Stable cost per execution does not necessarily mean stable cost to produce successful outcomes.
That gives us a different economic question to investigate.
What did one successful AI outcome actually cost?
Cost per Execution and Cost per Successful Outcome Answer Different Questions
Cost per execution is useful.
If an AI workflow consumes $100 of attributable cost across 100 executions, an average of $1.00 per execution tells us something meaningful about the cost of its runtime activity.
That can help us reason about execution behavior, compare versions of a workflow, investigate infrastructure or model changes, and understand how expensive it is to run the system.
The problem begins only when we ask that metric to answer a different question.
Suppose two workflows both cost:
$1.00 per executionbut one produces a successful outcome in 90 of 100 executions while the other produces one in 50.
The execution-level metric remains the same. The relationship between economic work and successful results does not.
| Question | Meaning |
|---|---|
| EXECUTION QUESTION | How expensive is the activity? |
| OUTCOME QUESTION | How much attributable cost corresponds to the outcomes classified as successful? |
One does not replace the other.
The first keeps the denominator close to runtime activity. The second relates cost to a result we have decided counts as successful.
That distinction matters especially in AI systems because a single result may involve variable model usage, retrieval, tool calls, external APIs, validation, retries, or other runtime work—and because not every execution necessarily produces the outcome the product cares about.
So a workflow can become cheaper to execute without necessarily becoming cheaper at producing successful results.
The reverse can also happen. Execution cost can remain stable while the number of successful outcomes changes.
Execution cost measures activity. Outcome cost relates economic work to successful results.
But that second view creates two problems that the arithmetic alone cannot solve.
What exactly belongs in the cost?
And what exactly qualifies as a successful outcome?
Those questions sit on opposite sides of a deceptively simple formula.
The Formula Is Simple. The Numerator and Denominator Are Not.
At its simplest, the relationship looks like this:
ATTRIBUTABLE COST
─────────────────────────
SUCCESSFUL OUTCOMESThe division is easy.
The definitions underneath it are not.
The numerator requires us to decide which economic work belongs inside the analysis. The denominator requires us to decide which outcomes qualify as successful.
If either side changes meaning, the resulting number changes meaning with it.
We will use cost per successful outcome descriptively for this relationship.
Related ideas already appear in AI systems and benchmarks as cost per successful task, cost per resolution, and cost per outcome. But those terms do not imply one universal definition of either cost or success.
That distinction matters.
A formula can produce a number with two decimal places while still hiding disagreement about what was measured.
The formula is simple. The numerator and denominator are not.
So before interpreting the result, we need to open both sides of the fraction.
Start with the numerator.
The Numerator Is an Attribution Decision
Imagine an AI coding workflow producing Outcome #123.
One outcome can involve more than one economic operation:
- Initial execution
- Model inference
- Retrieval
- Tool call
- Validation
- FAILED
- Retry
- Additional model call
- Validation
- PASSED
- Successful outcome
Now ask a seemingly simple question:
What did Outcome #123 cost?
Model, retrieval, tools, validation, failed paths, and shared infrastructure may all be relevant. Simply dividing every company-level AI expense by successful outcomes would not automatically produce a meaningful numerator.
For this Guide, attributable cost means the cost included within the analytical boundary chosen for the population of outcomes being measured. That boundary can include directly observed execution cost, defensible allocations of shared cost, and estimates; other costs may remain unknown or intentionally excluded. Those are not equivalent forms of evidence, and the metric should not pretend otherwise.
Consider two calculations:
Analysis A
Model inference
+ retrieval
+ external toolsand:
Analysis B
Model inference
+ retrieval
+ external tools
+ allocated infrastructure
+ human reviewBoth can be valid for a clearly defined purpose, but they are not automatically comparable merely because both are labeled:
cost per successful outcomeTheir numerators describe different cost boundaries. Total AI spend is not automatically the right numerator merely because it is easiest to obtain; the population above the line has to make analytical sense relative to the population below it.
And even after that boundary is defined, another complication remains.
Some of the economic work associated with a successful outcome may not itself have succeeded.
The Successful Attempt Is Not Necessarily the Whole Economic Story
Suppose Outcome #123 required two attempts.
Attempt 1
FAILED
$0.16
Attempt 2
SUCCESS
$0.18The successful attempt cost $0.18. But that does not automatically answer a different question:
What did it economically take to produce the successful outcome?
If the second attempt exists because the first one failed, it may be tempting to write:
$0.16 + $0.18
=
$0.34 per successful outcomeThat arithmetic is correct.
The attribution is not automatically established.
Depending on the analytical question, the first attempt may be attributable to the eventual outcome, treated as reliability overhead, allocated differently, or have an uncertain relationship to it.
So we can say:
Successful attempt cost
=
$0.18But before saying:
Cost of producing the successful outcome
=
$0.34we need an attribution decision.
The cost of the successful attempt is not necessarily the cost of producing the successful outcome.
The important point is not that every failed attempt belongs to the successful outcome that followed it. It is that a failed path can remain economically relevant while adding nothing to the successful-outcome count.
Failed work can affect the numerator without increasing the denominator.
That is one reason two systems with the same number of successful outcomes can require different amounts of economic work to produce them.
But solving the numerator is only half of the problem.
Even if we could perfectly reconstruct every relevant unit of cost, we would still need to decide which outcomes are allowed below the line.
The Denominator Is a Success Decision
Suppose an AI system records:
| Measure | Count |
|---|---|
| Executions | 1,000 |
| Completed executions | 850 |
| Outputs passing validation | 700 |
| Outcomes satisfying the chosen success condition | 600 |
Which number belongs in the denominator?
There is no universal answer.
We could calculate:
cost per executioncost per completed executioncost per validated outputcost per successful outcome
Each ratio answers a different question.
If we want to understand the economics of execution activity, executions may be the relevant denominator. If we want to understand the economics of outputs that passed a specific validation rule, validated outputs may be appropriate.
If we want to understand the economic work associated with outcomes that satisfied a defined success condition, then successful outcomes become relevant.
The last denominator is not automatically superior. It is simply aligned with a different question.
The denominator should match the economic question.
But once we choose successful outcomes, we inherit another requirement.
We need to know what successful means.
For one coding agent, success might mean:
required tests passedFor another analysis, it might mean:
change accepted into the codebaseFor another:
reported production problem resolvedThose conditions are not interchangeable.
And the evidence required to support them may not arrive from the same place or at the same time.
A deterministic validation can establish one kind of condition. A domain event can establish another. Human review, user behavior, or a model-based evaluation may provide evidence for others.
The purpose here is not to choose one universal evidence hierarchy.
It is to make the denominator explicit.
The denominator is only as defensible as the rule and evidence used to classify outcomes as successful.
A number called successful outcomes can look precise while still hiding a changing definition underneath it.
And there is another complication.
Sometimes we do know what success means.
We simply do not know the answer yet.
Pending Is Not Failed
Suppose 100 AI executions have completed.
At the moment we inspect them, their outcome state looks like this:
| Measure | Count |
|---|---|
| Executions | 100 |
| SUCCESSFUL | 70 |
| FAILED | 10 |
| PENDING | 20 |
Assume the attributable cost for those executions is $100.
If we divide that cost by the 70 outcomes currently confirmed as successful, we get:
$100 / 70
≈ $1.43 per confirmed successful outcomeThat calculation may be correct for the information currently available.
But it does not mean the other 30 outcomes all failed.
Ten are classified as failed. Twenty are still pending.
That distinction matters when success depends on evidence that arrives after execution.
A coding agent may complete before a review is approved. A support workflow may respond before the customer confirms whether the issue was resolved. An AI-generated recommendation may be produced before the downstream application event that determines whether the intended outcome occurred.
Now imagine that seven days later the same population looks like this:
| State | DAY 1 | DAY 7 |
|---|---|---|
| SUCCESSFUL | 70 | 85 |
| FAILED | 10 | 10 |
| PENDING | 20 | 5 |
No new execution needs to have occurred for the denominator to change.
What changed was the available outcome evidence.
Using the same illustrative $100 cost boundary:
DAY 1
$100 / 70
≈ $1.43 per confirmed successful outcome
DAY 7
$100 / 85
≈ $1.18 per confirmed successful outcomeThe apparent economics changed even though the execution population and the cost assigned to it did not.
The denominator can change because evidence arrived, not because new execution occurred.
This is why PENDING should not automatically be collapsed into FAILED.
And it is why the observation window matters.
If the success condition depends on delayed evidence, a recent cohort may contain outcomes whose final classification is not yet known.
That does not make the metric unusable. It means the metric needs context.
A cost-per-success figure based only on currently confirmed outcomes describes what is known at that point in time.
It should not silently pretend that unresolved outcomes have already reached their final state.
Cheap Execution Does Not Necessarily Mean Cheap Successful Work
We can now return to the two workflows from the beginning.
| Measure | Workflow A | Workflow B |
|---|---|---|
| Executions | 100 | 100 |
| Attributable cost | $100 | $100 |
| Successful outcomes | 90 | 50 |
| Cost per execution | $1.00 | $1.00 |
| Cost per successful outcome | $1.11 | $2.00 |
Their execution-cost profiles are identical in this simplified example.
Their observed outcome yield is not.
Workflow A produced 90 outcomes classified as successful from the same 100 executions and $100 of attributable cost. Workflow B produced 50.
That difference becomes invisible if we look only at cost per execution.
This does not make cost per execution wrong. It means cost per execution is answering the question it was built to answer.
It tells us about the economics of execution activity. It does not incorporate how often that activity produces the result we have chosen to count as successful.
Cheap execution and cheap successful work are different properties.
This distinction is already appearing in contemporary AI-agent benchmarking.
A 2026 benchmark published by Arize AI and Fireworks evaluated 2,400 agent runs across 10 models and 40 tasks using, among other measures, cost per successful task. Its methodology calculates that measure from spend across attempts relative to successful runs, making success rate and attempt cost jointly relevant to the comparison.
That can change how systems compare.
A model or agent that is inexpensive per attempt can still require more economic work per successful task if it succeeds less often or requires additional attempts.
Conversely, a more expensive execution can sometimes correspond to a lower cost per successful task if it produces successful results more reliably.
The important lesson is not that one metric should replace the other.
It is that the comparison can depend on the question.
| Question | Meaning |
|---|---|
| EXECUTION QUESTION | How expensive is the activity? |
| OUTCOME QUESTION | How much economic work corresponds to the successful results? |
Those questions become especially important in agentic systems because the path to one outcome may vary substantially from another.
One execution may complete in a single model call. Another may require retrieval, multiple tools, validation, retries, or additional model calls.
Two systems can therefore look similar at one analytical boundary and different at another.
But even if we build a defensible measure of cost per successful outcome, we still have not answered the most important business question.
We know something about the cost of producing success.
We do not yet know what that success was worth.
Cost per Successful Outcome Is Not Value or Profitability
Suppose two AI workflows each produce 100 successful outcomes with $100 of attributable cost.
| Measure | Workflow A | Workflow B |
|---|---|---|
| Successful outcomes | 100 | 100 |
| Attributable cost | $100 | $100 |
| Cost per successful outcome | $1.00 | $1.00 |
At this analytical boundary, their cost per successful outcome is identical.
But that does not mean the outcomes are economically equivalent.
One workflow might produce a relatively low-value result. Another might complete work that customers value much more highly. Their commercial models may differ.
Their revenue may differ. And other relevant costs may sit outside the boundary used in this calculation.
The ratio does not tell us any of that.
COST PER SUCCESSFUL OUTCOME
≠
VALUE PER OUTCOMEA successful outcome is not automatically a unit of customer value.
It is an outcome that satisfied the success condition chosen for the analysis.
That condition might be highly correlated with customer value. Or it might represent only one intermediate step toward it.
For example, an AI coding agent might successfully produce a change that passes the required tests.
That is a meaningful success condition.
But it does not, by itself, tell us how valuable the change is to the customer, how much revenue it supports, or whether it ultimately improves the business outcome the customer cares about.
This distinction also prevents another tempting shortcut:
LOW COST PER SUCCESSFUL OUTCOME
≠
PROFITABLEImagine:
| Measure | Workflow A | Workflow B |
|---|---|---|
| Cost per successful outcome | $0.50 | $4.00 |
| Revenue associated with it | $0.30 | $20.00 |
These numbers are deliberately simplified, and they are not a complete profitability calculation.
But they expose the problem.
Workflow A has the lower cost per successful outcome. That alone does not make it economically better.
Workflow B has a much higher cost per successful outcome, yet the commercial context around that outcome is also very different.
Profitability requires more information than the cost side of the ratio can provide.
Depending on the question, that may include revenue, commercial terms, additional cost categories, customer or workload mix, and the economic boundary being analyzed.
So there are at least three different questions hiding near each other:
| Question | Meaning |
|---|---|
| RUNTIME QUESTION | What does execution activity cost? |
| OUTCOME QUESTION | What does it cost to produce results classified as successful? |
| BUSINESS QUESTION | What are those successful outcomes worth in their commercial and economic context? |
These are not a universal hierarchy.
They are different analytical questions.
A system can improve on one while remaining unchanged—or even moving in the opposite direction—on another.
Cost per successful outcome tells us something about the cost side of producing success. It does not tell us what that success was worth.
And keeping those questions separate matters for another reason.
Measuring economics around an outcome does not mean the outcome has to become the thing the customer buys.
Measuring per Outcome Does Not Mean Charging per Outcome
Suppose an AI product internally measures:
$1.20 attributable cost
per successful outcomeNothing about that measurement determines how the product must be sold.
The commercial model could still be:
Subscription
Credits
Usage-based pricing
Seats
Outcome-based pricing
or a combination of theseThe commercial model can still be a subscription, credits, usage-based pricing, seats, outcome-based pricing, or a combination. Real products already use concepts such as resolutions or outcomes in operational reporting and, in some cases, commercial measurement.
For example, Intercom defines multiple Fin AI Agent outcome types and uses resolutions within its commercial model. Zendesk similarly uses automated resolutions as a measure of AI-agent usage.
But the existence of outcome-based commercial models does not turn cost per successful outcome into a pricing prescription.
These are separate decisions.
- INTERNAL ECONOMIC VIEW
- How much attributable cost corresponds to successful outcomes?
- ≠
- COMMERCIAL MODEL
- What does the customer buy, consume, or get charged for?
A company can sell a subscription, deduct credits, charge for usage, or use an outcome-based commercial unit while separately analyzing the cost of each successful outcome.
Measuring economics per successful outcome does not require selling the product per successful outcome.
Customers do not need to absorb the complexity of the runtime for the business to measure it. Conversely, an outcome-based price does not automatically tell the business what that outcome cost to produce.
Suppose a customer is charged:
$5 per resolved taskThat commercial fact does not reveal whether the attributable cost behind one resolution was:
$0.40or:
$4.50or whether the cost varies substantially across different kinds of tasks.
Commercial measurement and economic measurement can intersect without becoming the same thing.
That brings us to a final danger.
Once a metric feels closer to the outcome the product cares about, it is easy to treat it as the number that finally tells us whether the system is improving.
It does not.
A better denominator can still produce a metric we misunderstand.
Don't Turn It Into Another Vanity Metric
Suppose your dashboard shows this:
Cost per successful outcome
August $1.80
September $1.20
Change -33%At first glance, that looks like an improvement.
Perhaps it is.
But the metric itself does not tell us why it changed.
The underlying execution may have become cheaper. The success rate may have improved. The workflow may require fewer retries. A provider rate may have changed.
The mix of workloads may be different. The attribution boundary may have changed. The success condition may have changed.
Or the observation window may contain a different proportion of outcomes whose success is still pending.
Several of these changes can happen at the same time.
So even after moving from execution-level cost toward outcome-level economics, an old analytical problem remains:
A change in cost per successful outcome is an observation before it is an explanation.
Consider two very different scenarios.
| Scenario | Attributable cost | Successful outcomes | Cost per successful outcome |
|---|---|---|---|
| Scenario A | $180 → $120 | 100 → 100 | $1.80 → $1.20 |
| Scenario B | $180 → $180 | 100 → 150 | $1.80 → $1.20 |
Here, the measured improvement comes from the numerator.
Now consider Scenario B.
The resulting metric is identical.
The mechanism underneath it is not.
In Scenario A, less attributable cost corresponds to the same number of successful outcomes. In Scenario B, the same attributable cost corresponds to more successful outcomes.
Those are economically different stories hidden behind the same final ratio.
And there are more possibilities.
A lower provider rate could reduce the numerator without changing runtime behavior. A model change could increase the number of successful outcomes while also increasing cost per execution.
A retry-policy change could alter both cost and success yield. A shift toward easier workloads could improve the aggregate metric even if the economics of each workload remained unchanged.
Or nothing about the underlying system may have improved at all.
The measurement itself may have changed.
Comparisons Require Stable Meaning
Suppose that in August an outcome counted as successful when:
validation passedbut in September the rule became:
user accepted the resultNow compare:
August $1.80 per successful outcome
September $1.20 per successful outcomeThe arithmetic still works.
The comparison may not mean what it appears to mean.
The denominator no longer represents the same success condition.
The same problem can occur above the line.
Perhaps August includes only directly observed model and tool costs, while September also includes allocated infrastructure. Or perhaps the valuation basis changed because the applicable provider rates changed.
Or one period gives outcomes seven days to resolve while another measures them after only twenty-four hours.
When comparing cost per successful outcome across periods, workflows, models, or customers, interpretation therefore depends on knowing whether important parts of the measurement remained comparable.
That includes, among other things:
cost boundary
success definition
valuation basis
observation windowThese do not have to remain frozen forever.
Definitions may change when the analytical question changes or the measurement improves.
But those changes need to be visible.
Otherwise, a measurement change can masquerade as an economic change.
A metric can change because the system changed—or because the meaning of the metric changed.
Even a consistently defined aggregate can hide another problem.
| Level | Cost per successful outcome |
|---|---|
| Aggregate | $1.40 |
| Workflow A | $0.80 |
| Workflow B | $1.20 |
| Workflow C | $4.70 |
The aggregate can be useful.
It is also incomplete.
A change in workload mix could move the aggregate even if none of the individual workflows became more or less efficient at producing successful outcomes.
Aggregate outcome economics can hide materially different workflow economics.
That does not mean every metric needs to be segmented endlessly.
It means the level of aggregation should match the question being investigated.
The same principle has followed us throughout this Guide.
The number is not the explanation.
A More Outcome-Relevant Denominator Does Not Remove the Need for Explanation
Cost per execution can hide differences in success yield.
Cost per successful outcome can expose some of those differences.
That makes it useful.
It does not make it self-explanatory.
Imagine the metric rises next month:
$1.20 → $1.65What happened?
Perhaps executions became more expensive. Perhaps executions cost exactly the same, but fewer produced successful outcomes. Perhaps retries increased.
Perhaps the underlying resource rates changed. Perhaps the workload mix changed. Perhaps more outcomes remained pending at measurement time.
Perhaps the attribution boundary changed. Perhaps the definition of success became stricter.
Looking only at the final ratio cannot distinguish among those explanations.
To understand the change, we still need evidence from underneath the metric.
Execution evidence.
Cost evidence.
Attribution evidence.
Outcome evidence.
And commercial context when the question extends beyond cost.
A more outcome-relevant denominator does not remove the need for explanation.
This matters because improving the metric without understanding the mechanism can be just as misleading as optimizing the wrong metric.
A lower cost per successful outcome may be desirable.
But before treating it as evidence that the system became economically better, we need to know what produced the change.
The metric tells us that the measured relationship changed.
Attribution can help us locate where the change occurred.
The underlying evidence helps us explain what changed.
What Did Success Actually Cost?
We started with two workflows that looked identical from one perspective.
| Measure | Workflow A | Workflow B |
|---|---|---|
| Executions | 100 | 100 |
| Attributable cost | $100 | $100 |
| Cost per execution | $1.00 | $1.00 |
Then we added the outcome:
| Measure | Workflow A | Workflow B |
|---|---|---|
| Successful outcomes | 90 | 50 |
| Cost per successful outcome | $1.11 | $2.00 |
The arithmetic revealed something that cost per execution could not.
But the arithmetic was never the difficult part.
To interpret the numerator, we needed to know what economic work belonged inside the boundary. To interpret the denominator, we needed to know what counted as success and what evidence supported that classification.
Failed attempts could consume resources without becoming successful outcomes. Retries could increase the economic work associated with reaching one. Pending outcomes could leave the denominator temporarily incomplete.
And even after all of that, cost per successful outcome still could not tell us what the outcome was worth, whether the product was profitable, or how the customer should be charged.
Those are different questions.
The useful distinction is therefore not:
| Label | Metric |
|---|---|
| BAD METRIC | cost per execution |
| GOOD METRIC | cost per successful outcome |
It is:
- RUNTIME ACTIVITY
- What did execution cost?
- different question
- SUCCESSFUL RESULTS
- What economic work did it take to produce the outcomes we count as successful?
Both views can matter.
They simply illuminate different boundaries of the system.
For AI products, that distinction becomes particularly useful because the path between execution and outcome can contain variable model usage, retrieval, tools, external services, validation, retries, and uncertain or delayed success evidence.
The same commercial product can therefore produce the same number of executions while requiring very different amounts of economic work to produce successful results.
And the same cost per successful outcome can emerge from very different mechanisms underneath.
So the goal is not to find one number that finally summarizes AI economics.
It is to make the economic question explicit enough that the number has a defensible meaning.
The formula is simple. The numerator and denominator are not.
And once both sides are defensible, one question remains:
If successful outcomes became more expensive, do you know whether the cost changed, the success rate changed—or your definition did?
References
Arize AI & Fireworks AI — Cost per Successful Task Benchmark — Benchmark of AI-agent economics across 2,400 runs, 10 models, and 40 tasks, including cost per successful task and the economic effect of attempts that do not successfully complete the benchmark task.
Arize AI — Fireworks Cost Benchmark Repository — Public benchmark methodology and implementation, including the calculation of cost per successful task from total spend across attempts and successful runs.
FinOps Foundation — Allocation — Framework for allocating technology costs, including directly attributable and shared costs and the use of allocation methodologies where costs cannot be mapped directly to a single unit.
Intercom — Fin AI Agent Outcomes — Documentation describing Fin outcome types, resolution rules, and the operational and commercial treatment of AI-agent outcomes.
Zendesk — About Automated Resolution Tiers — Documentation describing automated resolutions as a measure of AI-agent usage and their role in Zendesk's commercial model.