THE SIGNAL IN ONE SENTENCE

A failed AI agent leaves more behind than a red score. It leaves the page it saw, the command it chose, the tool response it received, the state it changed and the point where a better decision might have rescued the task. That trail is expensive evidence. The Agent Error Dataset asks whether we can stop throwing it away. The new preprint introduces a collection of 50,228 error-diagnosis pairs drawn from 9,961 source tasks across 33 environments, 19 agent-harness families and 23 policy models. The authors call the construction method Agentic Error-to-Training, or AET. It gathers natural failures, identifies a decision to revise, proposes a correction, checks the diagnosis against the recorded trace and, where the environment supports it, replays the alternative from the same checkpoint. The basic idea is excellent. Most teams treat an agent run as success or failure. Success enters the demo reel. Failure enters a bug tracker, perhaps accompanied by the technical equivalent of a shrug. A useful failure record is richer. It can show what information was available at the moment of error, whether the model ignored a tool result, whether the harness hid an instruction, whether the grader was broken, and whether a proposed fix actually changed the outcome. The paper reports a promising replay result. Across 3,062 matched pairs, the first proposed correction passed the verifier 51.1 percent of the time. Retrying the original action from the same checkpoint passed 18.4 percent of the time. That is a gain of 32.7 percentage points under matched execution settings. This is stronger evidence than asking a model whether its own correction sounds sensible. The alternative was executed. But execution is not absolution. A correction that passes one verifier in one restored environment does not prove the diagnosis found the unique root cause. It proves that the proposed branch worked under those conditions. Several different actions may rescue the same task. A grader may reward the wrong thing. A flaky environment may turn yesterday's fix into today's ghost story. The authors state this distinction directly, and it should be printed on every dashboard that turns agent failures into training data. The dataset is also easy to overcount if you read the largest number and keep walking. The 50,228 figure counts error-diagnosis pairs, not 50,228 independent agent runs. The current index connects those pairs to 38,278 stored source-trace blobs and 9,961 source tasks. Multiple policies, seeds, temperatures or proposed diagnoses can produce separate records around related underlying work. Collection membership does not mean a record is ready for every training objective. The paper therefore creates different views. Diagnosis training can use a trace-supported attribution even when a correction was not replayed. Recovery training requires an executed continuation that passes. Preference training requires comparable alternatives with the same relevant input. That separation is practical. A dataset should not grant every row the same evidentiary passport simply because all the rows live in one folder. The diagnosis-training result is interesting and more limited than the headline version. The authors fine-tuned Qwen3-8B on full diagnoses from 1,656 source tasks. On a 943-case holdout, exact-step agreement with recorded internal teacher labels rose from 47.2 percent for the base model to 63.6 percent, averaged across three seeds. The strongest prompted reference shown scored 54.7 percent. Mean agreement increased at four larger training-set sizes. That demonstrates that the model learned the recorded labeling system. It does not, by itself, demonstrate that the recorded labels are true. The paper is admirably blunt about this. The comparison measures agreement with teacher annotations, including their conventions and defects. A separate historical human audit did not relabel the 943-case test set. The teacher is part of the measurement apparatus, not a little oracle living above the chart. This matters because diagnosing a long agent run is not like checking whether two plus two equals four. A trace may contain several defensible error locations. The harness may omit the system prompt or tool list the original model saw. An infrastructure failure may masquerade as a model error. A reference answer may fail its own grader. In a planner-executor system, the label can even blame an agent that does not appear in the exported transcript. The authors found examples of exactly those problems. A case-by-case audit of an earlier frozen diagnosis release rejected 37 of 60 sampled rows. The paper says that release retained 54 rows from tasks whose reference answer failed the grader. It also says 125 of the 943 held-out rows were planner-executor transcripts whose gold responsible agent was a unit not present in the transcript. The current construction checks address these kinds of defects, but the paper says they do not retroactively validate the frozen inputs used for the main diagnosis experiment. That does not make the result worthless. It makes the result specific. The training improved agreement with a historical internal labeling process. That is useful evidence about learnability and scale. It is weaker evidence about general diagnostic truth. A factory can teach an inspector to match yesterday's stamps perfectly while yesterday's stamp pad contains mistakes. The human review provides another useful slice. Four paper authors, each with doctoral or AI research experience, completed an AI-assisted audit of 80 historical records. Every record received two judgments. When both reviewers selected a preferred error step, they chose the same step in 59 of 69 cases, or 85.5 percent. Verdict categories agreed in 50 of 80 cases. Both reviewers accepted 41 of 80 records. The sample was chosen for coverage, not as a random estimate of the final collection's quality, and the reviewers were paper authors using shared AI assistance. The result is informative, not a universal accuracy certificate. The public-transfer evidence is similarly mixed. On a public attribution benchmark, the three trained seeds improved over the base model under the paper's unified protocol. Once flagged task overlap was excluded, the interval included zero. On TrajErrBench, mean accuracy remained below the base model. The authors also note that undocumented outside exposure cannot be ruled out. This is the awkward but useful part of research: one table says the model learned, another asks what it learned, and a third wonders whether the exam questions were already in the room. Actor training is the most direct test of whether diagnoses help an agent do better rather than merely talk about mistakes. In a single-seed WebShop-lite comparison, action-only repair training scored 6.67 percentage points higher than success-only training. The result suggests corrected actions can provide useful supervision. It remains one seed, one recipe comparison and a text-based environment. The paper does not establish that failure-trained agents will reliably transfer to messy production systems, multimodal interfaces or high-consequence tools. The practical lesson is not to avoid failure data. It is to build a chain of evidence around it. Start by preserving the full state that the agent actually saw. Record the model and harness versions, system instructions, available tools, budgets, environment response, grader and terminal outcome. A summary written after the fact is not enough if it silently adds information the agent never had. Then separate three questions. Where did the run first become unrecoverably wrong? Why was that decision wrong given the evidence available then? What alternative action should be tested? Those questions often produce different answers. A late visible error may be a symptom of an earlier assumption. A plausible explanation may point at the right step for the wrong reason. A good correction may work without proving the diagnosis. Replay is the receipt. Restore the relevant checkpoint, keep the policy, harness, budget and verifier settings matched, and execute both the proposed correction and a fresh retry of the original action. Retain both outcomes. If the correction wins once, label it as supported under those conditions. If it wins repeatedly across seeds or perturbations, confidence can rise. If the environment cannot be restored, keep the record in a weaker evidence tier rather than pretending prose is execution. The verifier deserves its own audit. Teams should test whether reference answers pass, whether harmless alternatives are accepted, whether broken states can accidentally score, and whether the grader changes across versions. Some tasks need more than one verifier. A web purchase agent might need a task-completion check, a policy check, a spending check and a human review of consequence. Passing one narrow grader should not erase failure somewhere else. Teacher diversity also matters. If one powerful model writes the diagnosis, another model can challenge the error location, a deterministic rule can check trace citations and a human can review sampled disagreements. The goal is not consensus theater. It is to expose where the label depends on one model's habits. Disagreement can be a high-value training record, but only if it remains visible instead of being flattened into one confident gold label. Data splits should group related tasks, traces and revisions together. If one failed run lands in training while a near-duplicate diagnosis of the same source task lands in evaluation, the score can become a memory test. Public benchmark overlap, prior evaluation exposure and repeated environment templates need separate registries. When overlap cannot be eliminated, report the contaminated and cleaned results side by side. For teams operating agents today, the dataset suggests a better incident loop. Capture the failed trace automatically. Quarantine infrastructure and grader problems. Ask for a trace-cited diagnosis and a concrete alternative. Replay when possible. Route uncertain or high-consequence cases to a person. Store evidence grades. Train only on rows that meet the needs of that objective. Re-evaluate on task families, environments and tool arrangements the model did not see during training. This can improve products without waiting for a giant research release. A coding agent can learn that running a targeted test before editing is better than guessing. A support agent can learn that an account lookup must precede a policy promise. A browser agent can learn that a disabled button is evidence, not an invitation to invent a workaround. A data agent can learn that a failed query may indicate a schema mismatch rather than missing customer activity. Each lesson is stronger when the corrected action is executed and the outcome is recorded. Privacy and security still apply. Failed traces can contain credentials, customer records, proprietary code and the exact exploit path an agent attempted. A failure dataset needs redaction, access controls, retention limits and a policy for which fields can enter training. Debugging value is not a permission slip. The most informative trace may also be the one most dangerous to copy into a broad model pipeline. The Agent Error Dataset is valuable partly because its paper keeps tripping over its own measurement problems and writes them down. That is not a weakness to hide. It is the subject. Turning failures into training data is possible, and matched replay provides a credible way to test corrections. Turning model-written diagnoses into truth is harder. The paper's audits, overlap checks and frozen-release defects show how quickly a useful pipeline can teach a model to imitate a flawed teacher. The plain signal is simple: failed agent runs are not garbage, but they are not labels either. They are evidence waiting for an investigation. Preserve the trace, test the correction, audit the grader, keep the teacher's fallibility visible and let uncertainty survive long enough to be useful. Otherwise the system does not learn from its mistakes. It learns to describe them in the house style.

01

WHAT ACTUALLY CHANGED

The Agent Error Dataset was submitted as a preprint on September 30, 2026.

The collection contains 50,228 error-diagnosis pairs linked to 9,961 source tasks.

The collection spans 33 text-based environments, 19 harness families and 23 policy models.

The current index connects the pairs to 38,278 stored source-trace blobs.

The five-stage AET pipeline collects failures, produces diagnoses, checks trace support, replays corrections where possible and creates objective-specific training views.

Across 3,062 matched replay pairs, first corrections passed at 51.1 percent versus 18.4 percent for original-action retries.

Full-diagnosis fine-tuning used 1,656 source tasks and a 943-case holdout.

Qwen3-8B exact-step agreement with internal teacher labels rose from 47.2 to 63.6 percent across three seeds.

The strongest prompted reference shown scored 54.7 percent on the same internal-label comparison.

A single-seed WebShop-lite comparison favored action-only repair training over success-only training by 6.67 percentage points.

The paper reports a separate AI-assisted human audit of 80 historical records.

The paper documents defects in an earlier frozen diagnosis release and says newer checks do not retroactively validate the main experiment inputs.

02

WHY THIS MATTERS

Failed agent runs contain observations, actions and environment responses that a final reward discards.

Executed replay is stronger evidence than a model saying its proposed fix sounds correct.

A passing correction supports one alternative under defined conditions without proving a unique root cause.

Error-diagnosis pairs are not the same thing as independent agent executions.

Different training objectives require different evidence thresholds.

Agreement with internal teacher labels measures learnability and convention matching, not universal diagnostic truth.

Broken graders, missing tool context and infrastructure faults can turn bad records into confident supervision.

Task overlap and prior evaluation exposure can exaggerate apparent transfer.

A useful failure pipeline needs provenance for models, harnesses, tools, graders and environment state.

Human review is most valuable when it targets disagreement and high-consequence records.

Failure traces can contain sensitive data, credentials and exploitable system details.

Keeping evidence grades visible helps prevent uncertain records from becoming fake ground truth.

FIG. 275TURN A FAILED RUN INTO EVIDENCE
1Preserve the exact trace→
2Locate the decision to revise→
3Explain with visible evidence→
4Propose one concrete alternative→
5Replay from the same checkpoint→
6Audit the verifier→
7Grade the evidence→
8Train and test on new settings
A failure becomes useful supervision only after the trace, correction, execution result and grader survive separate checks.

03

WHERE IT COULD HELP

  • Capture the exact state, instructions, tools and observations available before the failed action.
  • Keep model, harness, environment, grader and budget versions with every failure record.
  • Separate error location, causal explanation and proposed correction into distinct fields.
  • Require diagnoses to cite evidence that was visible to the acting model.
  • Replay a proposed correction from the same checkpoint whenever the environment supports it.
  • Compare the correction with a fresh retry of the original action under matched settings.
  • Label a passing correction as supported under tested conditions rather than uniquely correct.
  • Audit whether reference answers and harmless alternatives pass the verifier.
  • Use independent judges and deterministic checks to surface label disagreement.
  • Route uncertain, disputed and high-consequence records to trained human reviewers.
  • Group related tasks, traces and diagnoses together when splitting training and evaluation data.
  • Track public benchmark overlap and prior evaluation exposure separately.
  • Train diagnosis, recovery and preference objectives only on rows that meet their evidence requirements.
  • Evaluate transfer across unseen task families, harnesses, tools and environments.
  • Redact secrets and personal data before failure traces enter a training pipeline.
  • Version the dataset so corrected labels do not silently rewrite old evaluation results.

KEEP A HAND ON THE WHEEL

This article covers a September 30 preprint produced by the dataset authors. The 50,228 figure counts error-diagnosis pairs, not independent executions or training-ready examples. Replay results use a 3,062-pair cohort and support corrections only under the tested conditions. The diagnosis model is measured mainly against recorded internal teacher labels, including known historical defects, rather than a fully human-relabeled ground truth. The WebShop-lite actor comparison uses one seed. The public-transfer results are sensitive to task overlap, and mean TrajErrBench accuracy remained below the base model. The 80-record human audit was coverage-selected, AI-assisted and conducted by paper authors, so it does not estimate final-collection accuracy.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 1, 2026.

PUBLICATION RECEIPT: Revision 1. Published October 1, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US