AI Agent Benchmarks Miss a 24-Point Reliability Gap
An IBM Research post on Hugging Face says ALTK-Evolve consistency guidelines lifted AppWorld Pass^5 results from 53.0% to 69.0% while mean accuracy also improved.

Hugging Face carried an IBM Research post that puts a reliability problem behind AI-agent benchmark scores: an agent can look accurate on average while failing to repeat the same successful run.
According to IBM Research, in the AppWorld test, five repeats with a GPT-4.1-backed ReAct agent produced a 77.4% mean success rate, while only 53.0% of tasks cleared all five attempts, a 24.4-point gap between average capability and repeatable performance.
Mean@k measures average pass rate across repeated runs.
Pass^k measures the share of tasks where every run succeeds.
A customer-support workflow, contract checker or transaction-reconciliation agent does not only need a high average; it needs the same request to produce the same correct path when the task has not changed.
IBM's explanation centers on decision stability inside an agent trajectory.
At each step, the model chooses an action, argument, retry or tool call from a distribution of possible next tokens.
A sharp distribution has an obvious winner and tends to repeat.
A flatter distribution can leave several options close together, making small endpoint-level differences enough to reorder the winning choice.
Because an agent chains many decisions, a small chance of a flip at each step can accumulate into a materially different run.
The post says that problem is not removed by greedy decoding or a fixed seed.
The ReAct setup ran at temperature 0.0, so the variance was not ordinary sampling.
Hosted endpoints can still produce slight probability shifts, and near-tied choices can resolve differently even when the prompt, model and task stay the same.
ALTK-Evolve's new consistency-guideline flow tries to find those unstable points before they break a future run.
Its Consistency Analyzer replays each recorded decision step through controlled resampling, using one additional model call per step and drawing five completions by default.
The process works against the already recorded context rather than making new tool calls or replaying the full task, and it produces a scorecard that identifies which decisions are most likely to flip.
Those flagged steps then become targeted guidelines in the existing ALTK-Evolve format.
In one AppWorld example, the system generated instructions to count checkbox-style markers with a line-anchored regex rather than a plain substring count, and to verify search results by checking for multiple note matches before proceeding.
The point was not to memorize a single answer, but to turn a fragile decision pattern into a reusable operating rule.
The reported results show a larger change in repeatability than in average accuracy.
Across AppWorld test_normal's 168 tasks, consistency guidelines raised aggregate Pass^5 from 53.0% to 69.0%, while Mean@5 rose from 77.4% to 81.0%.
That narrowed the consistency gap from 24.4 percentage points to 12.0 points, and nearly a third of previously inconsistent tasks became tasks the agent passed on every run.
The gains were concentrated where the benchmark had more room to improve.
Medium tasks rose 22.9 percentage points in Pass^5, a 44% relative gain, while hard tasks rose 14.3 points, a 45% relative gain.
Easy tasks gained 12.2 points.
Mean@5 held or improved at every difficulty level, which the post presents as a safeguard against merely trading average accuracy for repeatability.
Transfer tests suggest the guidelines were not limited to the exact trajectory that produced them.
Applied to a related task in the same AppWorld scenario, the same-task improvement was only three points higher than the similar-task result.
IBM Research reports that with the weaker gpt-oss-120b model, same-task Pass^5 rose from 10.1% to 16.1%, while similar-task generalization improved by 8.7 points.
The practical takeaway is a measurement one: agent teams may need to publish consistency beside accuracy when they move systems into workflows where reproducibility is part of the product.
The open-source ALTK-Evolve repository has added the analyzer and guideline-generation pieces from the experiments, so developers can examine fragile decisions that a headline average would otherwise mask.




















