THE SIGNAL IN ONE SENTENCE
An AI agent can fail on step forty-seven because it misunderstood step twelve, papered over step twenty-three and confidently blamed the final tool. That is why debugging agents is miserable. The logs are long. Several agents may be talking. Tools return noisy results. One mistake changes the context for everything after it. By the time the workflow collapses, the visible error may be a symptom with three generations of bad decisions behind it. AgentRx, an open-source diagnostic framework from Microsoft Research and academic collaborators, tries to make that mess auditable. It takes a failed execution trace, normalizes the events into a common format, generates constraints from tool schemas, policies and the trace itself, checks those constraints step by step, and records each violation with supporting evidence. An LLM judge then uses that validation log to identify the first unrecoverable failure and assign a root-cause category. The plain signal is simple: do not ask a model to stare at a wall of agent logs and improvise a postmortem. Give the judge a structured trail of claims it can inspect, dispute and trace back to the run. Microsoft highlighted AgentRx on October 2 alongside the project's public code, paper and benchmark data. The repository uses the MIT license. The benchmark and paper use CC BY 4.0. There is an important version wrinkle. Microsoft's research page summarizes a benchmark of 115 failed trajectories. The current paper, revised August 31 and linked from that page, reports 170 trajectories across eleven task settings. The public repository also describes the expanded framework and links the current benchmark. The 170 figure is the one used here because it comes from the latest primary paper. This is not a trivial editorial footnote. Debugging tools should make provenance visible, and their own release materials should do the same. The expanded benchmark combines four groups of failed runs. Thirty-nine come from Tau-bench retail workflows, where one agent uses structured APIs to modify orders, process returns and answer customer questions. Forty-two come from Flash, a multi-agent system for diagnosing cloud incidents. Forty-four come from Magentic-One, where an orchestrator coordinates browsing, file and coding agents on open-ended tasks. The remaining forty-five are drawn from nine settings in two earlier failure datasets, including software development, API workflows, mathematical reasoning and web or file tasks. Each trajectory contains messages, tool calls, tool outputs and observable environment state. Human annotators marked every failure they could find, then identified the critical one: the earliest failure in the causal chain from which the system did not recover. That definition matters. The first mistake is not always the root cause. An agent can issue a bad search, notice the problem and repair it. The last error is not necessarily the root cause either. A failed upload may be inevitable after an earlier agent invented a filename that never existed. AgentRx aims for the first unrecoverable step, the place where the run crossed from messy into doomed. The benchmark was expensive to label. Three annotators used a grounded-theory process, allowing categories to emerge from the traces before freezing the taxonomy. On a thirty-trajectory agreement sample, the paper reports a mean pairwise kappa of 0.89. Across all 170 traces, annotation took 62.5 human hours. The median Tau-bench trace had thirty-six steps. Magentic-One had a median of thirty-three and a range reaching 130. Annotators averaged twenty to twenty-four minutes per trajectory depending on the domain. That is the labor AgentRx hopes to reduce. The framework begins by converting different log formats into a shared trajectory representation. That sounds dull because it is dull. It is also foundational. A retail API call, a cloud troubleshooting sub-agent and a web-browsing action do not naturally describe themselves in the same vocabulary. After normalization, AgentRx builds two kinds of constraints. Global constraints come from the tool schemas and any available domain policy. They can express rules such as a required field, an allowed value, a confirmation requirement or the relationship between a tool response and the agent's claim. Dynamic constraints are generated for each step from the task and everything observed up to that point. They capture facts that only become relevant during the run. If a product lookup returns ten options, a later agent should not report eleven. If a user has not supplied the color needed for an exchange, the agent should not invent one. The constraints have guards that decide when they apply and assertions that return satisfied or violated. Some checks run directly over structured data. Others are semantic checks evaluated by a language model. When a check fails, AgentRx writes the step, the violated constraint and the evidence into a validation log. Only then does the final judge enter. The judge sees the original task, the trajectory, a checklist for the failure taxonomy and the evidence-linked violation log. It predicts the critical step and a category, then provides a rationale. The taxonomy contains nine substantive root-cause categories: failure to follow instructions or a plan, invention of new information, invalid tool invocation, misinterpretation of tool output, intent and plan misalignment, underspecified user intent, unsupported intent, triggered guardrails and system failure. The implementation also allows an inconclusive result when the evidence is insufficient. This is a better-shaped question than "what went wrong?" For every category, the checklist asks concrete questions. Was required information actually missing? Did the tool report a schema error? Did the agent's reasoning contradict the tool output? Was the requested action impossible with the available tools? Did an access policy, login or safety guardrail block execution? Those distinctions stop one convenient label from swallowing every failure. An invalid tool invocation needs a different fix from an unsupported request. A misunderstood user goal needs a different fix from a network outage. Calling all of them hallucinations is not diagnosis. It is shrugging with a technical accent. The published results are promising and incomplete. The paper says AgentRx improved step localization by 75 percent on average over prior work and improved root-cause categorization by 30 percent over the authors' judge baselines. Those are relative summary figures across experiments, not a claim that the system is 75 percent accurate. Exact performance varied sharply by domain. In the paper's comparison against a modified Who and When baseline, AgentRx located the exact critical step in 42.7 percent of Tau-bench traces, 36.4 percent of Magentic-One traces, 82.5 percent of Flash traces and 61.5 percent of the forty-five RelWork traces. Root-cause category accuracy ranged from 46.2 percent to 65.9 percent across those groups. That range is the reality check. On long open-ended traces, the system often gets the exact step wrong. A 36.4 percent exact score on Magentic-One is useful evidence of improvement over the comparison baseline, not a replacement for an experienced engineer. Relaxing the metric makes the tool look more practically helpful. On Magentic-One, 59.8 percent of predictions landed within five steps of the human label. On RelWork, 92.6 percent landed within five. A diagnosis that points an engineer to the right neighborhood can save time even when it misses the exact line. This is one reason the validation log may matter more than the headline label. If the tool says the critical failure was step thirty-seven, an engineer should be able to inspect the evidence at steps thirty-four through forty and decide whether the story holds. A naked root-cause label asks for trust. An evidence log invites review. Step-by-step constraint generation helped most on longer traces. The alternative is to read the full trajectory once and generate all constraints in one shot. That is cheaper. It is also more vulnerable to context dilution, where details buried in a long run lose influence. For Magentic-One, the paper reports that step-by-step evidence reduced the average distance from the annotated failure from 22.5 steps to 12.4 in one configuration. Short Flash traces showed a much smaller difference. The accuracy has a price. Under the paper's GPT-5 pricing assumptions, a judge-only baseline cost about five cents per trajectory. The one-shot AgentRx configuration cost eighteen cents, roughly 3.6 times more. Constraint generation accounted for about 81 percent of that cost. Those are experimental estimates, not a current cloud invoice. They reveal the tradeoff: structured evidence costs more than one-pass opinion, and detailed step-by-step analysis costs more than generating constraints once. A sensible production design would not run the most expensive mode on every successful trace. Use cheap monitoring to identify suspicious runs. Preserve the full trace. Trigger one-shot diagnosis for ordinary failures. Reserve step-by-step analysis for high-impact incidents, ambiguous chains or workflows whose earlier mistakes are expensive to miss. The framework is a postmortem tool, not a seatbelt. AgentRx diagnoses a failed run after the trajectory exists. It does not prove that an agent is safe to deploy. It does not prevent the agent from changing a customer record, leaking a secret or deleting a file. It does not repair the workflow automatically. The same constraint machinery could eventually support runtime checks, but that is not the claim being evaluated here. There is another agent inside the diagnostic loop. The default experiments use GPT-5 for constraint generation and judging. The paper also tests o3 and DeepSeek-V3.2, showing that the pipeline is not tied to one model family. Results with DeepSeek improved in some domains and were neutral or worse in others. That means the diagnosis inherits model behavior. Semantic checkers can misunderstand the trace. Generated constraints can be weak, redundant or wrong. The final judge can overweight a visible violation that is merely a downstream symptom. The paper's limitations section says false positives and noisy validation signals can misdirect the judge. This is not an embarrassing exception. It is the central reason to keep the evidence visible. An AI-generated audit log is not ground truth. It is a structured hypothesis about the run. The strongest checks are often the boring deterministic ones. Did the tool call include every required field? Did the value match the schema? Did the agent state a number different from the tool output? Did the final record contain an action that policy forbids? These checks can be evaluated directly. Semantic questions are fuzzier. Did the agent understand the user's real intent? Was a missing detail necessary to proceed? Did a later recovery genuinely repair the earlier mistake? Those calls need context and judgment. Teams should label the difference. Every violation in an internal AgentRx-style system should say whether it came from executable code, a model-based semantic check or a human rule. Reviewers should know which evidence is hard and which evidence is interpretive. Logs also carry risk. A useful agent trajectory may contain customer messages, account identifiers, document contents, authentication failures, internal tool names, policy text and sensitive side effects. Centralizing that material for diagnosis creates a valuable debugging record and an equally valuable breach target. The paper warns that trajectory data may contain sensitive information. A production deployment needs redaction before analysis, access controls, short retention where possible, encrypted storage and an audit trail for the audit trail. It also needs stable versions. Store the agent model, system prompt, tool schemas, policy, code commit, environment state and diagnostic model version with every incident. Otherwise a later rerun may produce a different explanation and no one will know which component changed. The public dataset has a small access wrinkle of its own. Its card is publicly visible under CC BY 4.0, but downloading the files requires agreeing to share contact information. The visible card currently lists Tau retail and Magentic-One splits, while the paper's 170-trajectory benchmark also includes Flash and RelWork. That does not invalidate the paper. It does mean independent users should verify which benchmark portions and annotations they can actually retrieve before claiming a full reproduction. AgentRx is immediately useful as a design pattern even for teams that never run its code. Normalize agent events into one trace. Write explicit invariants. Separate global policy from facts discovered during the task. Evaluate constraints at the step where they become relevant. Attach evidence to every diagnostic claim. Keep a fixed taxonomy so incident counts mean the same thing across weeks. Let people override the model and record why. Then connect diagnosis back to prevention. If invalid invocations dominate, improve schemas and argument validation. If agents misread tool output, add typed parsers and read-back checks. If missing user information causes failure, force a clarification before state changes. If guardrails repeatedly block valid work, fix permission design rather than teaching the agent to push harder. If system failures dominate, stop tuning the prompt and repair the infrastructure. The point of a postmortem is not to produce a beautiful explanation. It is to change the system that failed. AgentRx does not solve agent reliability. It offers something more practical: a way to turn a sprawling trace into claims that engineers can inspect, challenge and count. That is a worthwhile shift. The model still gets a vote. The evidence gets the record.
01
WHAT ACTUALLY CHANGED
Microsoft Research published the AgentRx framework, code and linked benchmark materials for diagnosing failed AI-agent trajectories.
The current paper expands the benchmark to 170 failed trajectories across eleven task settings, superseding the 115-trajectory count on Microsoft's shorter research page.
AgentRx normalizes heterogeneous logs, generates global and step-specific constraints, checks them and records violations with supporting evidence.
An LLM judge uses the validation log and a fixed taxonomy checklist to identify the first unrecoverable failure and assign a root-cause category.
The repository implements a six-stage pipeline from trajectory normalization through invariant generation, checking, judging and report plots.
The code is MIT licensed, while the benchmark and paper use CC BY 4.0.
02
WHY THIS MATTERS
The final visible error in a long agent run may be a downstream symptom rather than the decision that made recovery impossible.
Evidence-linked violations are easier to inspect and contest than a model-generated root-cause label with no audit trail.
A shared taxonomy lets teams count whether failures come from plans, invented facts, tool calls, output interpretation, missing information, guardrails or infrastructure.
Step-by-step constraints can preserve local context that gets diluted when a judge reads one long trajectory all at once.
The published exact-step scores show meaningful improvement while remaining too uneven for unsupervised incident decisions.
Diagnosis can guide concrete fixes to schemas, prompts, permissions, parsers, infrastructure and human approval gates.
03
WHERE IT COULD HELP
- Generate evidence-linked postmortems for failed customer-service, coding, browsing and incident-response agents.
- Normalize messages, tool calls, outputs and environment changes into one versioned trajectory format.
- Create deterministic checks for schema errors, impossible values, missing confirmations and contradictions with tool output.
- Use semantic checks for intent and reasoning questions, while labeling them as model judgments.
- Route high-impact or ambiguous failures to step-by-step diagnosis and use cheaper one-shot checks for routine incidents.
- Track root-cause categories over time so reliability work targets recurring failure patterns.
- Link each diagnostic claim to the exact trace window and policy or tool rule that supports it.
- Record human overrides of the diagnostic judge and feed disagreement into evaluator improvement.
- Redact credentials, personal data and document contents before centralizing agent traces.
- Store the agent, prompt, tool, policy, code and diagnostic versions with every incident.
KEEP A HAND ON THE WHEEL
AgentRx diagnoses failed traces after the fact and does not prevent unsafe actions, prove deployment safety or automatically repair an agent. The current paper reports 170 trajectories, while Microsoft's research summary still says 115, so readers should use the versioned paper when comparing counts. Exact critical-step accuracy ranges widely across the four evaluated groups and remains low on long open-ended Magentic-One traces. The framework still uses language models for semantic checks, constraint generation and final adjudication; false positives, weak constraints and downstream symptoms can misdirect the judge. Experimental costs rely on the paper's GPT-5 pricing assumptions and are not a current production quote. The visible Hugging Face card is public but file access requires sharing contact information, and its listed splits do not obviously expose every group described in the expanded paper. Trajectory logs may contain sensitive user, policy, tool and system data. Watch for independent reproductions of the full 170-trajectory benchmark, measured engineer time savings, privacy-preserving logging, stronger deterministic checks, calibration of inconclusive results and evidence that diagnosis leads to fewer repeated production failures.
04
TERMS WORTH KEEPING
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on October 4, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US