THE SIGNAL IN ONE SENTENCE

An AI can provide a convincing explanation for a decision without accurately identifying the factors that actually changed its behavior.

01

WHAT ACTUALLY CHANGED

Researchers tested eight models from the Claude, GPT, and Gemini families in two synthetic settings: an advisor making recommendations and a monitoring system classifying prompts for risk or harm. Each model made a decision and named the three factors it claimed had most influenced that answer.

The researchers then treated those explanations as hypotheses. They changed the cited factors to test whether each one was necessary, and they retained individual factors while removing others to test whether each one was sufficient to preserve the decision.

The claimed rankings aligned only weakly or moderately with measured influence. Mean Spearman correlations were 0.349 for necessity and 0.354 for sufficiency in the advisor task. In prompt monitoring, the corresponding figures were 0.431 and 0.580.

An uncited factor influenced the advisor more than its lowest-ranked cited factor in roughly 58 percent of cases. The gap was smaller in prompt monitoring, but it remained large enough to undermine the idea that a tidy explanation automatically constitutes an audit trail.

02

WHY THIS MATTERS

A model-generated rationale is easy to mistake for direct access to the decision process. It may instead be a plausible story assembled after the answer exists. The difference matters when explanations are shown to doctors, moderators, auditors, applicants, or operators deciding whether to trust the system.

The behavioral test offers a more disciplined approach. If the model says a factor mattered, changing that factor should predictably affect the output. If the decision barely moves, the explanation deserves less confidence, no matter how polished its prose appears.

This does not make explanations useless. It changes their status from evidence to a claim that can be checked. Apparently the AI press secretary also needs a fact-checker.

FIG. 062MAKE THE EXPLANATION PROVE IT
1MODEL DECIDES→
2MODEL NAMES FACTOR→
3CHANGE FACTOR→
4MEASURE DECISION→
5COMPARE CLAIM
A stated reason becomes more credible when changing it produces the behavioral effect the explanation predicts.

03

WHERE IT COULD HELP

  • Test explanations used in automated risk reviews
  • Perturb recommendation inputs to measure actual influence
  • Audit high-stakes rationales before operators rely on them
  • Escalate decisions when explanations fail behavioral checks

KEEP A HAND ON THE WHEEL

The study examines individual decisions in two synthetic environments. It does not establish that every explanation is false or that the measured perturbations capture every real causal pathway. Production audits should combine behavioral tests with domain expertise and system-level evidence.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 7, 2026.

PUBLICATION RECEIPT: Revision 1. Approved by Zak and published September 7, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US