THE SIGNAL IN ONE SENTENCE
An AI can provide a convincing explanation for a decision without accurately identifying the factors that actually changed its behavior.
01
WHAT ACTUALLY CHANGED
Researchers tested eight models from the Claude, GPT, and Gemini families in two synthetic settings: an advisor making recommendations and a monitoring system classifying prompts for risk or harm. Each model made a decision and named the three factors it claimed had most influenced that answer.
The researchers then treated those explanations as hypotheses. They changed the cited factors to test whether each one was necessary, and they retained individual factors while removing others to test whether each one was sufficient to preserve the decision.
The claimed rankings aligned only weakly or moderately with measured influence. Mean Spearman correlations were 0.349 for necessity and 0.354 for sufficiency in the advisor task. In prompt monitoring, the corresponding figures were 0.431 and 0.580.
An uncited factor influenced the advisor more than its lowest-ranked cited factor in roughly 58 percent of cases. The gap was smaller in prompt monitoring, but it remained large enough to undermine the idea that a tidy explanation automatically constitutes an audit trail.
02
WHY THIS MATTERS
A model-generated rationale is easy to mistake for direct access to the decision process. It may instead be a plausible story assembled after the answer exists. The difference matters when explanations are shown to doctors, moderators, auditors, applicants, or operators deciding whether to trust the system.
The behavioral test offers a more disciplined approach. If the model says a factor mattered, changing that factor should predictably affect the output. If the decision barely moves, the explanation deserves less confidence, no matter how polished its prose appears.
This does not make explanations useless. It changes their status from evidence to a claim that can be checked. Apparently the AI press secretary also needs a fact-checker.
03
WHERE IT COULD HELP
- Test explanations used in automated risk reviews
- Perturb recommendation inputs to measure actual influence
- Audit high-stakes rationales before operators rely on them
- Escalate decisions when explanations fail behavioral checks
KEEP A HAND ON THE WHEEL
The study examines individual decisions in two synthetic environments. It does not establish that every explanation is false or that the measured perturbations capture every real causal pathway. Production audits should combine behavioral tests with domain expertise and system-level evidence.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
Grounding
Connecting an AI answer to specific outside information that can support it.
OPEN GLOSSARY CARD
Benchmark
A fixed test used to compare how systems perform on the same tasks.
OPEN GLOSSARY CARD
Evaluation gate
A required test or review that a system must pass before it advances to the next stage.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 7, 2026.
PUBLICATION RECEIPT: Revision 1. Approved by Zak and published September 7, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US