THE SIGNAL IN ONE SENTENCE
When an AI reports a molecular measurement to three significant figures, a low error may mean it remembers the published number rather than understands the chemistry well enough to predict it.
01
WHAT ACTUALLY CHANGED
Researchers at Hamburg University of Technology, Helmholtz-Zentrum Hereon, the German Research Center for Artificial Intelligence, and Saarland University audited 22 frontier language models across 12 molecular regression benchmarks. The preprint was submitted September 4 and appeared in the newest public research listing on September 7.
The usual score asks how close a model comes to the benchmark value. That cannot reveal whether the model inferred a property from molecular structure or retrieved a number encountered during training. The researchers instead examined whether answers reproduced the second and third significant digits more often than a molecule-blind predictor could achieve from the distribution of labels alone.
Retrieval was widespread but concentrated. On five of the 12 datasets, more than half of the tested models showed statistically significant verbatim retrieval. The remaining datasets produced only isolated flagged cases. One model reached a median absolute error of 0.025 kilocalories per mole on FreeSolv, far below the 0.6 kilocalories per mole default experimental uncertainty assigned to most entries.
Reasoning made the retrieval easier to see. With the same molecules and prompts, 47 of 264 model-and-benchmark combinations were flagged at each endpoint's minimum reasoning setting. At the roughly 1,024-token setting, 89 combinations were flagged, an 89 percent increase.
The team also tried to interrupt retrieval by scrambling molecular structure strings and hiding the original property name while providing 100 examples. Flagged cases fell from ten of 12 combinations to three of 12. That experiment weakened the chemistry along with the identifying clues, so it demonstrates that retrieval can be interrupted, not that the damaged benchmark remains a valid chemistry test.
02
WHY THIS MATTERS
Molecular-property benchmarks influence how people judge models for chemistry, materials science, toxicology, and drug research. If one model remembers more of the answer sheet, a leaderboard can reward training-data exposure while appearing to measure scientific prediction.
The third significant digit is the useful tell. A capable chemist can estimate whether a molecule is broadly soluble or toxic from its structure. Predicting the last digit of an experimental result is different because that digit often sits inside measurement noise. Reproducing it repeatedly begins to look less like chemistry and more like recognition.
The reasoning result is especially awkward. Turning the effort setting down can make a benchmark appear cleaner even though the same model recovers the values when allowed to think longer. A contamination audit therefore needs to report the reasoning setting and treat a low-effort result as a lower bound.
This does not mean every molecular answer is memorized or that language models cannot learn general chemical patterns. The researchers found clean and untestable cases too. Their sharper claim is that accuracy alone cannot separate prediction from retrieval, and benchmark designers need controls that can.
03
WHERE IT COULD HELP
- Audit molecular benchmarks before comparing frontier models
- Report reasoning settings alongside scientific evaluation scores
- Prefer recent or less widely redistributed benchmark collections
- Use digit-level tests to distinguish prediction from exact retrieval
- Run blinded controls when a benchmark shows contamination
KEEP A HAND ON THE WHEEL
This is a preprint and the detector is designed to find strong digit-level retrieval, not every form of contamination. A clean result is weak evidence when too few answers reach three significant figures. The blinding method destroys useful chemical information, and the study does not establish the score a strong classical chemistry model would receive under the same detector.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
SMILES
A text notation that represents a molecule's atoms, bonds, and structure.
OPEN GLOSSARY CARD
Benchmark contamination
When evaluation material appears in a model's training data and can make its test score look better than its general ability.
OPEN GLOSSARY CARD
Significant figures
The meaningful digits used to express the precision of a measured or calculated value.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 8, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 8, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US