THE SIGNAL IN ONE SENTENCE

An AI can summarize a paper convincingly and still fail the harder research job of finding the exact evidence, checking it across formats, and returning something another scientist can verify.

01

WHAT ACTUALLY CHANGED

Researchers led by the Chinese University of Hong Kong and Shanghai Artificial Intelligence Laboratory introduced SciDocBench, a benchmark designed around the work a scientific research assistant actually performs. The preprint was submitted September 4 and appeared in the newest public research listing on September 7.

The benchmark contains 124 expert-authored questions spanning five scientific domains, seven capability groups, and 19 subtasks. The tasks include recovering reading order from complicated pages, locating the figure or table supporting a claim, extracting notation and experimental details, checking numerical consistency, combining results across papers, tracing datasets, and connecting a paper with its code.

Every underlying question appears in four matched forms. The question is written in either English or Chinese, and the source documents arrive either as page images grouped first or as extracted text and visual elements interleaved in reading order. That creates 496 evaluation instances, but the authors correctly warn that these are four versions of 124 problems, not 496 independent scientific questions.

The strongest evaluated system, Claude Opus 5, scored 62.6 out of 100. No model led more than two of the seven capability groups. The paper reports persistent weaknesses in document perception, structured information extraction, cross-document synthesis, evidence localization, and robustness to how the same material is presented.

The researchers also describe SciDocIR, a typed representation that preserves pages, blocks, layout, equations, tables, cross-references, and provenance. They used it to construct roughly 15,000 supervised fine-tuning examples and 8,000 reinforcement-learning examples across 14 verifiable subtasks. The repository linked by the paper returned a not-found response during our September 8 verification, so the public paper is available but the promised project files could not yet be independently inspected.

02

WHY THIS MATTERS

Scientific reading is a relay race across formats. A claim may appear in the prose, depend on an equation, be supported by one panel of a figure, use a dataset defined in another paper, and rely on code stored elsewhere. Testing each skill separately can miss whether the baton survives the handoffs.

A polished summary is therefore a dangerously generous test. The useful assistant must show which source supports the answer, preserve the exact measurement or notation, distinguish evidence from interpretation, and produce a result that another person or program can inspect. Research does not become trustworthy because the paragraph sounds like it owns a lab coat.

The presentation result is also practical. For 11 of the 14 evaluated models, the performance gap between the two document representations was larger than the gap between English and Chinese questions. The same scientific content can become easier or harder depending on whether the system sees a page, extracted Markdown, or a different ordering of the evidence.

The benchmark points toward a more useful product standard. A research assistant should not receive one vague accuracy score. It should be tested separately on locating evidence, extracting structure, checking claims, joining documents, reconstructing data, and tracing the path from a paper to its code and dataset.

FIG. 070FOLLOW THE CLAIM BACK TO THE EVIDENCE
1READ FULL PAPERS→
2LOCATE OBJECTS→
3CHECK RELATIONS→
4JOIN SOURCES→
5RETURN RECEIPTS
A useful scientific assistant connects prose, equations, figures, tables, code, and datasets while preserving the route another researcher can verify.

03

WHERE IT COULD HELP

  • Evaluate literature-review assistants on complete scientific workflows
  • Require page, figure, table, equation, or code provenance for answers
  • Compare page-image and interleaved-document input pipelines
  • Train models on tasks with deterministic or executable checks
  • Audit whether extracted results remain reusable by researchers and software

KEEP A HAND ON THE WHEEL

This is a public preprint built from 124 underlying questions and 116 source papers. Computer science, mathematics, and selected natural and biomedical sciences receive more coverage than social sciences, humanities, and several engineering fields. Cross-document tasks use two to four papers rather than dozens, some scoring relies on model judges, and errors in layout parsing, source alignment, or OCR can propagate into the training data. The project repository linked by the paper was unavailable when verified.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 8, 2026.

PUBLICATION RECEIPT: Revision 1. Published September 8, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US