TRACES Benchmark Tests AI Discovery By Auditing The Investigation, Not Just The Answer
Apodex’s TRACES benchmark evaluates AI systems across tool use, repair, alternatives, coherence, evidence and scope, using 20 environments drawn from a larger registry of high-value research problems.

Apodex describes TRACES as a benchmark for examining the methods AI systems use during discovery work, alongside their final answers.
An iTNews Asia interview with Apodex AI research scientist Brian Wang frames the benchmark as a response to a gap in conventional AI evaluation.
General benchmarks usually ask a model to solve a known task, then score the final output.
TRACES instead places an AI system inside a working environment and examines whether it can choose tools, gather evidence, test alternatives and change course when the investigation demands it.
The name sets out the six behaviours Apodex wants to measure: Tools, Repair, Alternatives, Coherence, Evidence and Scope.
Together, those categories check whether a system used the right instruments, recovered from errors, considered competing explanations, stayed consistent across a long chain of work, tied claims back to evidence and made clear what its conclusion does not prove.
The source describes TRACES as a benchmark testbed for AI teams building agents, tool-use systems, solver frameworks and advanced models.
Those developers need diagnostic evidence about whether a system can sustain a complex investigation across many steps, not merely whether it can generate a plausible paragraph or pass a short coding task.
Apodex also positions the tool for research groups, life-sciences companies and industrial enterprises that hold difficult data or high-value questions of their own.
In those settings, a fluent but wrong answer can be expensive, especially when physical-world feedback arrives slowly or a technical report must be audited before it supports a business decision.
The benchmark’s problem set came from Apodex’s two-month mapping exercise.
The scan covered 561 industries grouped under 16 sectors and produced a registry of 423 high-value problems where progress can be limited by incomplete evidence, slow feedback loops, costly experiments, specialist data or multiple credible explanations.
From that registry, Apodex selected and developed the first 20 TRACES environments.
The initial set covers deployment-oriented intelligence and scientific engineering alongside clinical translation, biomedical discovery and frontier-model development.
Each environment is meant to turn a meaningful problem into a setting where an AI has to plan and investigate rather than retrieve a known answer.
The operating record is central to the benchmark.
TRACES assessments are anchored to recorded actions, clear scoring criteria and independent review where needed, so developers can inspect how a system reached a conclusion and where it failed.
Wang describes task-specific checks and repair loops for each environment.




















