An Agent That Succeeds 77% of the Time Is Not 77% Reliable
Agent evaluations report an average success rate. Running the same task five times tells a different story, and that second number is the one your business process actually depends on.
Antonio J. del Águila
Knaisoma
The pilot goes well. An agent handles the refund requests, the invoice reconciliation or the contract review, the team runs a hundred cases, and the deck says 77% success. Somebody asks whether that is good enough, the room decides it is, and the rollout starts. Three weeks later the support queue fills with cases the agent handled correctly in the pilot and got wrong in production, on inputs nobody changed.
Nothing regressed. The pilot measured the wrong thing. It reported how often the agent succeeded per run, averaged over the whole task set. The business process does not consume an average. It sends the same kind of request again next Tuesday, and again the Tuesday after that, and what it needs to know is whether the agent handles that request every time.
Two numbers, one benchmark
The distinction has a name and a formal definition, and it has been in the literature for two years. The team behind τ-bench, a benchmark for tool-using agents in retail and airline customer service, introduced pass^k in June 2024: the chance that all k independent trials of a task succeed, averaged across tasks. Its familiar sibling pass@k asks whether at least one of k trials succeeds. Both are estimated from the same repeated runs, and they answer opposite questions. The first is the reliability question. The second is the discovery question, and it is the right one only when a human or a checker will pick the good result out of the batch.
Their measurements set the scale of the problem. Using function calling, gpt-4o reached about 61% on τ-retail and about 35% on τ-airline at k equal to one. On τ-retail, the chance of solving a task on all eight of eight trials fell to roughly 25%.
A more recent result says the same thing with a stronger model and a different benchmark. IBM Research reported in September 2026 that a ReAct agent on GPT-4.1, evaluated on 168 AppWorld tasks, succeeded on 77.4% of runs while succeeding on all five repeated runs for only 53.0% of tasks. They call the 24.4 point difference the consistency gap, and they report that it reaches about 30 points on the hardest tier. The technical report carries the method and the full evaluation.
Two benchmarks, two model generations, two institutions, the same shape. The headline number that vendors and pilots report is an average per run. The number that predicts whether a repeated business process works is much lower, and nothing on the slide tells you by how much.
Why identical inputs produce different runs
Engineers reasonably assume this is a sampling setting they can turn off. It is not, for three separate reasons, and only one of them is under your control.
The first is the inference endpoint itself. Thinking Machines Lab published a careful analysis in September 2025 showing that temperature zero does not make a hosted endpoint deterministic, and that the usual explanation of concurrency plus floating point arithmetic is incomplete. The real cause is that inference kernels are not batch invariant: the same request produces slightly different arithmetic depending on how many other requests share its batch, and batch size varies with server load. Sampling 1000 completions at temperature zero from one open model gave them 80 distinct completions, with the first divergence at token 103. Batch-invariant kernels removed the variation entirely, which is a genuine fix, but it is a property of the serving stack. If you call a commercial API, you are not choosing it.
The second is the environment. A tool call returns a slightly different search ranking, a record was updated between runs, a rate limit fires on one attempt and not another. Agents amplify this because each observation feeds the next decision.
The third is the task description. Work on the reliability of computer-use agents separates execution stochasticity from ambiguity in the task specification and from behavioral variation across runs. Ambiguity is the one you can fix cheaply. If a request can be read two ways, the agent will sometimes read it each way, and no amount of model quality resolves that.
The gap is concentrated, which is the useful part
Here is an inference the published numbers support and neither paper spells out arithmetically. Suppose every task in a suite had the same per-run success probability and trials were independent. Then the chance of succeeding on all k trials would be that probability raised to the power k. For the AppWorld result, 0.774 to the fifth power is 27.8%. The measured value is 53.0%. For τ-retail, 0.61 to the eighth power is 1.9%. The measured value is about 25%.
The measured reliability is far higher than the uniform assumption predicts, and that is good news. It means outcomes within a task are strongly correlated: most tasks either succeed on every run or fail on every run, and the gap between the two numbers is carried by a minority of unstable tasks. IBM’s post describes the same structure in words, noting that roughly a quarter of the benchmark consists of tasks the agent can sometimes solve and sometimes cannot.
That changes what the fix looks like. If unreliability were spread evenly, you would need a better model across the board. Because it is concentrated, the work is triage: identify the unstable subset, and treat it differently from the part that already works every time. An average success rate cannot tell you which tasks those are. Only repeated runs of the same task can.
Pick k from the process, not from the benchmark
The benchmark chose k equal to five or eight for its own reasons. Your k comes from how the work arrives and what happens to the output. This is the question to settle before an acceptance criterion is written.
| How the work arrives | Metric that binds | Where effort pays |
|---|---|---|
| One-off exploration, a person reads every result and picks | pass@k at the number of samples you can afford | Generate more candidates. Extra sampling is the cheapest quality you will ever buy here. |
| Repeated task, output cheaply verifiable, retry costs only tokens | Average per-run success, plus verifier accuracy | Build the checker. A mediocre agent behind an exact verifier beats a better agent with no verifier. |
| Repeated task with an irreversible effect: a refund issued, an email sent, a record deleted | pass^k where k is how many times the process runs before anyone reviews the results | Put the side effect behind a confirmation step, or narrow the tool so the irreversible action needs an explicit approval. |
| Multi-step chain where each step consumes the previous output unchecked | The product of per-step success rates, which falls fast | Add checkpoints between steps so a failure is caught where it happens rather than at the end. |
| Mixed portfolio of task types, some stable and some not | Per-task pass^k, not the aggregate | Route the unstable subset to a human queue and keep the stable subset fully automated. |
The middle row is where most teams should be and few are. If the output can be checked programmatically and a retry is nearly free, run-to-run variation is an efficiency problem rather than a correctness problem. The row above and below it are where variance becomes a business risk, because a retry is either impossible or the error escapes unnoticed.
Two cautions on retries. A retry is only safe if the action is idempotent, and agent actions frequently are not: the first attempt may have already sent the message before failing on a later step. Recovering from that is a distributed systems problem, and we have argued elsewhere that treating a production agent as a distributed system is the durable answer. The second caution is that retrying until success turns your reliability metric into pass@k, which is the right metric only if you can tell success from failure without a human. If you cannot, retrying just gives you more chances to accept a wrong answer.
The cheapest improvement is usually not a bigger model
The IBM work is interesting less for the diagnosis than for what it cost to fix. Their Consistency Analyzer resamples decision points in an agent’s own recorded trajectory to find the steps where the model was close to choosing differently, then generates guidelines that stabilize those steps. Applying them raised the all-five-runs figure from 53.0% to 69.0% while the average moved only from 77.4% to 81.0%. On a related but different task in the same scenario, the guidelines still lifted reliability by 13 points, so they were not merely memorizing one trajectory.
Read that split carefully. Average accuracy barely moved. What improved was the agent doing the same thing twice. That is a strong argument that a large part of the consistency gap is not a capability ceiling and does not need a model upgrade to address, which matters when the reflex response to disappointing agent performance is to buy the more expensive tier.
Hold the result loosely. It is a preprint from the team that also ships the tooling, on one benchmark, with one primary model, and a weaker model gained considerably less. The mechanism is plausible and the direction is corroborated by the independent τ-bench finding that consistency and capability come apart, but the size of the improvement on your workload is an open question until you measure it. Treat it as a technique worth trying before a procurement decision, not as a number to put in a business case.
What to ask for, and what to measure
Vendor evaluations and internal pilots both default to the average, so ask explicitly. Request the per-task outcomes across repeated runs rather than a summary score, on a task set drawn from your own work, and ask how many runs the number represents. A supplier who cannot produce per-task repeat results is telling you they have not measured reliability, whatever the headline says.
Inside your own program, the acceptance criterion should name k and the consequence. Something in the shape of “this workflow runs about forty times a week with no human review before the customer sees the result, so we require the same case to succeed on all five consecutive runs for at least 90% of our case library, and everything below that goes to the review queue” is a criterion an engineering team can work against. A target expressed as average accuracy is not, because it never says which failures are acceptable.
Then keep measuring after launch. Model endpoints change under you, tool APIs change, and a task set that was stable in July may not be in October. The unstable subset is the thing to watch, because it moves first.
When the average is the right number
None of this argues for treating every agent workload as safety critical. If a person reviews each output before it goes anywhere, run-to-run variation costs review time and nothing else, and the average is a fine planning number. For drafting, summarizing, search and idea generation, sampling more and picking the best is not a failure mode, it is the intended use. The cost of over-engineering reliability into a workflow whose output a human reads anyway is real: slower delivery, more infrastructure, and an evaluation harness nobody maintains.
The distinction to hold onto is whether an unreviewed wrong answer reaches something that matters. Where it does, the average is not evidence, and a pilot that reports one is not finished.
If you are deciding whether an agent workflow is ready for unattended operation, the measurement is the hard part, not the model. We help engineering teams build evaluation harnesses that run repeated cases against real task libraries, set acceptance criteria that match how the process actually runs, and design the verification and human review boundaries around the tasks that turn out to be unstable. Talk with us about evaluating your agent workflows.
Stay updated
Get insights on engineering transformation delivered to your inbox.
Newsletter coming soon.