Your AI Workflow Completed. Did It Actually Succeed?
A framework for defining successful AI outcomes through explicit success conditions, supporting evidence, and economic interpretation.
Your AI Workflow Completed. Did It Actually Succeed?
An AI coding agent receives a task. It generates a code change, runs the required test suite, and finishes without an execution error.
The runtime records:
Coding Agent Run #1842
| Observation | State |
|---|---|
| Agent execution | COMPLETED |
| Code generated | YES |
| Tests | PASSED |
Everything looks successful. But suppose we keep looking:
| Claim | State |
|---|---|
| PR accepted? | UNKNOWN |
| PR merged? | UNKNOWN |
| Bug fixed? | UNKNOWN |
Did the run succeed?
Maybe.
If success means generate code that passes the required tests, then the evidence may already be enough. If success means produce a change that is accepted into the codebase, we do not know yet. If success means fix the reported problem for users in production, we may be even further away from knowing. The same execution can therefore support very different answers depending on the question being asked.
Before asking whether the run succeeded, we need a more precise question:
Succeeded at what?
That distinction matters once successful outcomes become part of how we understand an AI product's economics. A count of completed executions is a runtime measurement. A count of successful outcomes requires us to decide what qualifies as successful. Those are not automatically the same thing.
Completion Is Evidence of Something
The problem is not that COMPLETED is meaningless. It may be a perfectly accurate statement about the execution. The workflow reached its completed state. The agent returned an output. The required process finished according to the runtime's own semantics.
Likewise, tests = PASSED can be strong evidence that the generated code satisfied the test suite that was actually run. The mistake begins when we silently expand one of those observations into a larger claim.
- Agent execution = COMPLETED
- therefore
- Intended outcome = SUCCESSFUL
That conclusion requires more than the runtime status alone. Consider three possible questions about the same coding-agent execution:
| Question | Relevant evidence / what it can establish |
|---|---|
| Did the agent finish its execution? | COMPLETED can answer this. |
| Did the generated code pass the required tests? | Test results can answer this. |
| Did the change solve the problem users were experiencing? | Neither fact necessarily answers this. |
None of the earlier evidence becomes false or unimportant simply because the last question remains unresolved. It answers a different question.
This is the distinction that matters:
A runtime status can establish something about execution without establishing every claim we may want to make about the outcome.
And that means success cannot be treated as a label whose meaning is automatically inherited from the runtime. Before we can decide whether an outcome succeeded, we first need to define what would have to be true for us to call it successful.
Before Measuring Success, Define the Success Condition
Return to Coding Agent Run #1842. We already know:
| Observation | State |
|---|---|
| Agent execution | COMPLETED |
| Code generated | YES |
| Tests | PASSED |
Now define success as:
SUCCESS CONDITION\n\nThe generated code passes\nthe required test suite.For that condition, the observed test result may be exactly the evidence we need. The outcome can reasonably count as successful under that definition.
Now change the condition:
SUCCESS CONDITION\n\nThe generated change is accepted\ninto the codebase.The execution has not changed. The generated code has not changed. The test results have not changed. But the evidence we have is no longer enough to establish the condition.
| Observation | State |
|---|---|
| Tests passed | YES |
| PR accepted | UNKNOWN |
That does not make the outcome FAILED. It makes the success condition unconfirmed with the evidence currently available.
Change the condition once more:
SUCCESS CONDITION\n\nThe reported problem no longer\naffects users in production.Even a merged pull request may not, by itself, establish that claim. The important point is not that one of these definitions is the correct definition of success. It is that they describe different things.
A coding product designed to generate test-passing code may legitimately use the first condition. A system measuring accepted engineering work may care about the second. A team studying whether AI-generated changes actually resolve production problems may need the third.
The same execution can satisfy one success condition while another remains unresolved. There is no contradiction. The questions are different. And until the condition is explicit, a metric called successful_outcomes can hide that difference behind a single number.
Success Is Also an Evidence Problem
Defining the condition solves only half of the problem. Once we know what must be true, we still need to ask:
What evidence shows that it was true?
This gives us a simple reasoning model:
- EXECUTION
- What happened?
- SUCCESS CONDITION
- What must be true?
- SUCCESS EVIDENCE
- What shows that it was true?
This is not a required runtime architecture or a sequence every AI system must implement. It is a way to separate three questions that are easy to collapse into one.
The execution gives us observations about what happened. The success condition defines the claim we want to evaluate. The success evidence tells us what supports that claim.
Consider a simple condition:
SUCCESS CONDITION\n\nThe generated output conforms\nto the required schema.A deterministic schema validator can provide highly relevant evidence for that condition:
Schema validation PASSEDNow consider a different condition:
SUCCESS CONDITION\n\nThe generated report is useful enough\nfor the customer to make a decision.The same schema validation result may still tell us something about the report's structure. But it does not establish that the report was useful enough for the customer to act on it. The evidence did not become weaker. The claim changed.
That distinction matters:
Evidence is not strong or weak in the abstract. Its relevance depends on the claim it is supposed to support.
This is why success cannot always be reduced to finding the “best” success signal. A test result, validator, model-based evaluation, human approval, user action, or later domain event may each provide useful evidence. What matters first is what we are trying to establish. Success is not only a definition problem. It is an evidence problem.
Success Evidence Can Take Different Forms
Not every success condition can be evaluated in the same way. Some conditions can be checked deterministically. Others may be evaluated by a model.
| Evidence form | Example observation |
|---|---|
| Deterministic check | Required tests passed = TRUE |
| Model-based grader | quality_score = 0.92 |
| Human review | review_status = APPROVED |
| Pull request | status = MERGED |
These are not four levels of success. They are different forms of evidence that may be relevant to different success conditions.
A deterministic check can be exactly the right evidence when the condition itself is deterministic. A human review may be appropriate when acceptance requires human judgment. A domain event can directly establish that a particular domain event occurred. A model-based grader can provide scalable evaluation when the property being assessed cannot be captured by a simple deterministic check. But the existence of an evaluation does not remove the need to understand what it actually measures.
Suppose a model-based grader returns:
quality_score = 0.92That number may be useful evidence about whatever quality criterion the grader was designed to evaluate. It does not automatically mean:
92% probability\nthat the customer outcome succeededNor does a high quality score automatically establish that a pull request was accepted, a support issue was resolved, or a customer acted on a generated report. Those are different claims.
The right question is therefore not:
Which type of evidence is best?
It is:
Which evidence actually supports the success condition we are trying to establish?
The Same Success Judgment Can Carry Different Stakes
Suppose a team uses the model-based quality score to study an AI workflow internally. Across two versions of the system:
| Version | Average quality score |
|---|---|
| Version A | 0.81 |
| Version B | 0.89 |
That evidence may be useful for investigating whether the measured quality dimension improved. Now imagine using the same evaluator for a different decision:
If quality_score >= 0.85\ncount outcome as successful\nfor economic analysisThat gives the same judgment a different consequence. Now change the consequence again:
If quality_score >= 0.85\nconsume the customer's\noutcome allowanceThe evaluator has not changed. The score has not changed. But what the system is doing with that judgment has.
This does not mean a model-based evaluation can never support economic analysis or commercial rules. Nor does it mean human review is always required when the stakes increase.
The point is narrower:
The evidence standard appropriate for one use of a success judgment may not automatically be appropriate for another.
Internal experimentation, profitability analysis, customer-facing reporting, and commercial charging can ask different questions of the same evidence. A product can deliberately choose the same success condition and evidence standard across several of them. But that should be a design decision, not an accidental consequence of whichever status or score happens to be available.
This is especially important when economics enters the picture. If an outcome is going to count as successful in an economic analysis, the system needs to know what that label means and what evidence supports it. If the same judgment is going to trigger a commercial consequence, that is a separate decision again. Economic analysis and commercial treatment may rely on the same success evidence. They do not become the same question because they do. Before a success signal can safely travel across those boundaries, we need to know exactly what claim it represents.
Success May Arrive After Execution
So far, the examples have made success look like something we can evaluate as soon as execution ends. That is not always possible.
Return to the coding agent. Suppose the timeline looks like this:
- 10:00
- Agent execution completed
- 10:05
- Required tests passed
- 14:30
- Human review approved
- 16:08
- Pull request merged
When did the outcome become successful? Once again, the answer depends on the condition.
If the condition is:
The generated code passes\nthe required test suite.the relevant evidence may be available at 10:05.
If the condition is:
The generated change is accepted\ninto the codebase.the relevant evidence may not arrive until 16:08. The underlying AI execution finished at 10:00 in both cases. What changed was the point at which the evidence required for the chosen success condition became available.
- EXECUTION COMPLETION TIME
- does not necessarily equal
- OUTCOME CONFIRMATION TIME
This creates an important temporal distinction. The evidence required to confirm an outcome may arrive after the execution that produced it. For some products, that delay may be seconds. For others, it may be hours, days, or depend on a later human or domain event.
That does not make those products incapable of reasoning about successful outcomes. It means their outcome evidence has a different lifecycle from their execution state. And sometimes that lifecycle does not end with the first positive signal.
Later Evidence Can Change the Judgment
Consider an AI support agent. The agent responds to the customer, and the support system records:
| Observation | State |
|---|---|
| Agent responded | YES |
| Ticket | SOLVED |
If the success condition is simply:
The ticket entered\nthe solved state.then the observed status may establish exactly that. But suppose the question we care about is broader:
The customer's issue\nremained resolved.Twenty-four hours later:
| Observation | State |
|---|---|
| Customer replied | |
| Ticket | REOPENED |
The earlier SOLVED state did not become imaginary. It happened.
What changed is what the later evidence allows us to conclude about the broader success condition. Later evidence can change how an earlier success judgment should be interpreted. This does not mean every successful outcome must remain provisional forever. Nor does it require every product to wait indefinitely before counting success. It means that the success condition should make clear what is being claimed—and that some claims can only be supported by evidence that develops over time.
A simple boolean can hide that temporal reality:
outcome.success = TRUEmay look definitive even when the underlying judgment was based on evidence available at a particular point in time. The boolean may still be useful. But its meaning depends on the condition, the evidence, and when that evidence was observed.
A Successful Outcome Is Not Every Kind of Value
There is one more boundary worth preserving. Suppose our coding-agent success condition is:
The generated change was accepted\ninto the codebase.The pull request is merged. For that condition, the outcome can count as successful. But the merge does not automatically establish that:
| Claim |
|---|
| developer productivity increased |
| customer satisfaction improved |
| revenue increased |
Those are different claims. They may depend on additional events, longer time horizons, or many causes beyond the AI execution itself. This is why successful outcome should not quietly expand until it means all customer or business value created by the system.
A product may deliberately define success close to execution. Another may use a later domain event. Another may study downstream business effects separately. None of those choices is universally correct. What matters is that the success condition remains clear enough that the resulting metric still has a defensible meaning. A successful outcome does not automatically prove every downstream form of customer or business value.
That distinction also keeps the idea of a Value Unit separate from the success of a particular instance. A product may sell or expose a recognizable unit of value—such as a generated report, a support resolution, or a coding task—while still needing evidence to determine whether a particular instance satisfied the success condition being analyzed.
The unit tells us what kind of thing we are talking about. The success condition tells us what must be true for this instance to count as successful. Those are related questions. They are not the same question.
Now the Economic Problem Becomes Visible
Suppose an AI product records this for August. And this for September:
| Period | Completed executions | Successful outcomes |
|---|---|---|
| August | 1,000 | 850 |
| September | 1,000 | 700 |
The execution volume appears unchanged. The successful-outcome count fell. That difference could matter economically. The same amount of completed execution may now correspond to fewer outcomes satisfying the success condition being analyzed.
But before interpreting the change, we need to know what those numbers actually mean. What counted as successful in August? What counted as successful in September? Was the same success condition used? Was it supported by the same kind of evidence? Was the evidence available for all outcomes being compared? Did the confirmation window change?
Without those answers, 850 and 700 may be numerically precise while still representing different underlying judgments. A denominator can be numerically precise and still be conceptually unstable. This becomes especially important when successful outcomes enter an economic metric.
Consider the apparently simple expression:
cost
────────────────────────────
successful outcomesThe arithmetic is easy. The denominator is not.
Before that ratio can tell us something useful, successful outcomes needs a defensible meaning. That does not mean every product needs perfect knowledge of ultimate customer value. It does not mean every outcome requires human verification. And it does not mean the success condition has to sit at the furthest observable business event. It means the denominator should represent the thing the analysis claims it represents.
If success means required tests passed, count that deliberately. If success means change accepted into the codebase, use evidence appropriate to that condition. If success depends on a later domain event, account for the fact that confirmation may arrive later.
The economic problem begins when those distinctions disappear behind one generic label: SUCCESS
Same Completion Count. Different Outcome Yield.
Now consider two customers using the same AI workflow.
| Customer | Completed executions | Successful outcomes |
|---|---|---|
| Customer A | 100 | 92 |
| Customer B | 100 | 61 |
Assume both are evaluated using the same success condition and a comparable evidence standard. The customers generated the same number of completed executions. They did not produce the same observed number of successful outcomes under that definition. That difference is economically interesting. It is not yet an economic conclusion.
We still do not know: how much execution work belonged to those outcomes; what resources that work consumed; which rates applied; how much revenue was associated with each customer; why the observed outcome yield differed.
Customer B is not automatically less profitable. Customer A is not automatically more efficient. But 100 completed executions is no longer enough to describe what happened.
This connects two sides of the same economic problem. On one side: How much work belonged to the outcomes we are analyzing? On the other: Which of those outcomes actually count as successful? Both questions have to become defensible before combining cost and success into a meaningful economic interpretation.
Before Cost per Successful Outcome, Define Successful
A workflow can complete correctly. Its output can pass validation. A model-based grader can assign a high score. A human can approve it. A later domain event can confirm that something happened. Each of those observations can be useful evidence.
But evidence is always evidence of something. That is why the path from execution to economic success is not simply:
- COMPLETED
- SUCCESSFUL
A more defensible reasoning model is:
- EXECUTION
- What happened?
- SUCCESS CONDITION
- What must be true?
- SUCCESS EVIDENCE
- What shows that it was true?
- ECONOMIC INTERPRETATION
- What does that successful outcome mean for the question we are analyzing?
The last step does not make success a commercial charge. It does not prescribe outcome-based pricing. And it does not require every AI product to measure ultimate business value. It simply prevents an economic system from treating a runtime label as if its economic meaning were self-evident.
Earlier, the coding-agent run looked simple:
| Observation | State |
|---|---|
| Agent execution | COMPLETED |
| Code generated | YES |
| Tests | PASSED |
The difficult question was not whether those observations were true. It was what they allowed us to conclude. That is the boundary an economic analysis of successful outcomes eventually has to cross.
Before asking what a successful outcome cost, we need to know which work belonged to it. Before putting successful outcomes in the denominator, we need to know what qualifies an outcome to enter it. And before trusting that qualification, we need to know what evidence supports it.
The arithmetic comes later. The definition comes first.
What evidence is enough for an outcome to count as successful?
References
Amazon Bedrock, Evaluate the performance of Amazon Bedrock resources — supports the distinction between different forms of evaluation evidence, including automatic/model-based evaluation and human-based evaluation.
Amazon Bedrock, Evaluate model performance using another LLM as a judge — supports model-based evaluation in which an evaluator model scores responses according to selected or custom metrics, reinforcing that a model-generated score is evidence relative to the criterion being evaluated rather than a universal probability of outcome success.
GitHub Docs, About pull requests — checks, reviews, and merge as distinct parts of the pull-request lifecycle.
GitHub Docs, Merging a pull request — repository requirements such as reviews and status checks before merge.
Zendesk Help, About the ticket lifecycle and ticket statuses — solved tickets can be reopened when the requester responds.