THE SIGNAL IN ONE SENTENCE

An AI agent can spend six days writing code, running experiments, calling tools, launching sub-agents and drafting a paper. Then somebody has to work out what it actually did. That second job is becoming awkwardly large. Traditional AI evaluations often end with a score. The model passed or failed. The answer was right or wrong. A short transcript made it possible for a researcher to read the path and check whether the score hid something strange. Long-running agents leave a different sort of evidence. Their transcripts can stretch across hundreds of pages and contain messages, tool calls, tool responses, errors, retries, human interventions, context compactions and work delegated to other agents. A final score says almost nothing about which approach the system tried, where it became confused, what it handed off or how the evaluation setup shaped the result. The UK AI Security Institute has released an open-source package called Transect to make that record easier to inspect. Transect takes an evaluation transcript, information about the task and a vocabulary of activities the evaluator wants to distinguish. It then places recorded events, token use, sub-agent activity and model-generated labels on a shared turn-by-turn timeline. Click a labeled stretch and the report can take the reviewer back to the transcript passage beneath it. That last step is the important one. Transect does not solve a difficult transcript by asking another language model for a confident summary and stapling the word audit to the result. Its design assumes that an LLM judge can also be wrong or misleading. The package keeps provenance and reliability information with its classifications so a person can inspect how a label was produced and where judges disagree. The haystack becomes navigable. It does not become infallible. AISI and Meridian Labs built the package on Inspect Scout. The public repository is licensed under MIT, accepts Inspect evaluation logs and OpenClaw telemetry exports, and requires Python 3.12 or later. A structural pass can extract recorded events without calling a judge model. Judged analysis requires the relevant provider software and credentials and can incur model charges. Users define a reusable specification for an evaluation family. That specification describes expected phases such as planning, implementation, experimentation and write-up, plus labels for likely sub-agent roles. A judge model then classifies stretches of the transcript using that vocabulary. This is useful, but it creates a quiet source of discretion. The evaluator chooses the categories. The descriptions shape the labels. The judge model, sampling choices and verification settings can change the result. A category can be too broad, too narrow or simply absent. The repository itself recommends beginning with the structural report, drafting the vocabulary from what is present, inspecting the reliability audit and iterating. That is good scientific hygiene. It is also a reminder that a clean timeline is an interpretation, not raw reality. AISI demonstrates Transect on an open-ended AI research evaluation that generated almost 13 million tokens. The combined view suggested that agents spent much of their effort on operational work and manuscript production, with little evidence of a sustained hypothesis-generation stage. That finding is interesting because a polished paper can hide a shallow research process. An evaluator who looks only at the artifact might see competent prose, charts and code. The timeline can reveal whether the system spent time generating and testing explanations or mostly assembled a deliverable around the first workable path. But even here, caution earns its chair at the table. The tool can show what the selected labels and recorded events reveal. It cannot prove that an unrecorded mental process occurred. It cannot turn a judge label into ground truth. It cannot establish that every important behavior fits the evaluator's vocabulary. Agreement between judges can increase confidence in consistency while all of them remain wrong in the same direction. Sub-agent labels deserve special care. Transect can classify a sub-agent from the instruction it received. That describes what the sub-agent was asked to do, not necessarily what it did. The difference is the whole plot. Imagine a parent agent delegates a literature review. The sub-agent spends most of its time fixing broken data, then returns a thin list of papers. A dashboard that labels the span literature review from the delegation prompt is useful for orientation. A reviewer who treats that label as observed behavior will misunderstand the run. Provenance makes the correction possible. The reviewer can open the source turns, inspect the tool calls and compare the instruction with the work. That is why Transect matters beyond one government research institute. Companies are beginning to let agents operate across code repositories, browsers, spreadsheets, contract systems and internal knowledge bases. Long agent sessions create the same review problem as long evaluations. A team may know the final task completed and still lack a practical account of which data the agent used, what alternatives it tried, which tools failed and why it made a consequential choice. The right lesson is not to pour every private work transcript into one cheerful dashboard. Agent logs can contain source code, credentials, personal data, customer information, medical details, unpublished research and prompts copied from restricted systems. Observability without access control becomes a very organized leak. Teams need scoped collection, retention limits, redaction, role-based access and a clear separation between operational telemetry and sensitive task content. A reviewer should see enough evidence to investigate the claim without automatically inheriting every secret the agent encountered. The tool also needs a question before it needs a chart. For safety evaluation, that question might be whether the agent noticed and worked around a restriction. For research, it might be whether the system generated competing hypotheses before settling on one. For a customer-service agent, it might be whether escalation followed uncertainty or only followed failure. For a coding agent, it might be whether tests changed after the implementation broke them. Start with the behavior that matters. Define it carefully. Run a structural pass. Sample the source transcript. Add judge labels only where they reduce review effort. Compare judges. Inspect disagreement. Record the configuration. Keep the human able to reach the original evidence. Then test the analysis itself. Give multiple reviewers the same transcript and specification. Measure whether they reach similar conclusions. Change the vocabulary and see which findings disappear. Hide known events in a controlled example and test whether the pipeline recovers them. Track false reassurance, not just false alarms. This is slower than declaring the dashboard self-explanatory. It is also what makes the dashboard scientific. Transect arrives under an open license with code, examples and a paper. That lets outside evaluators inspect the implementation, run it on their own logs and challenge the method. The release does not establish that the package scales cleanly to every agent system, that its labels transfer across domains or that its judges are accurate enough for high-stakes decisions. Those are questions for use, replication and criticism. The larger signal is simpler. As agents gain longer horizons, evaluation cannot stop at the final answer. Researchers need to preserve a navigable path from a claim about behavior back to the turns, tool calls and settings that support it. The automated judge can point at the suspicious stretch. The human still has to read the evidence.

01

WHAT ACTUALLY CHANGED

The UK AI Security Institute released Transect as an open-source Python package built on Inspect Scout for long-horizon agent transcript analysis

Transect aligns recorded events, token use, sub-agent activity and model-generated behavior labels on one turn-based timeline

Each classification carries provenance and reliability information so reviewers can inspect source passages and judge disagreement

A structural pass can run without a judge model, while judged analysis requires provider credentials and can create model charges

A demonstration on an AI research evaluation covered almost 13 million tokens and surfaced heavy operational and manuscript work with little sustained hypothesis generation

02

WHY THIS MATTERS

Long agent runs can produce evidence volumes that make manual reconstruction slow and inconsistent

A final task score does not reveal failed approaches, delegation, tool errors, interventions or the path to the result

Using an LLM to judge another model moves part of the reliability problem rather than eliminating it

Traceable labels can help humans focus review while preserving a route back to the underlying transcript

The same observability pattern can help operational teams, but sensitive logs require access control, redaction and retention limits

FIG. 338From a long agent run to a reviewable claim
1An agent run records messages, tool calls, responses, interventions and sub-agent launches→
2A structural pass extracts recorded events and places them on a turn-based timeline→
3The evaluator defines task context, activity categories and expected sub-agent roles→
4Selected judge models label spans and expose votes, disagreement and reliability signals→
5A reviewer follows any claim back to the exact source turns and checks the interpretation→
6The team records the configuration, exports the data and tests whether another reviewer agrees
Automation narrows the reading queue. Provenance and human review decide whether the resulting interpretation deserves trust.

03

WHERE IT COULD HELP

  • Trace a suspicious safety-evaluation result back to the messages, tool calls and interventions that support it
  • Compare how several agents divide time among planning, implementation, experimentation and write-up
  • Inspect whether a coding agent changed strategy after test failures or merely rewrote the test
  • Review how customer-service or business agents escalate uncertainty, delegate work and recover from tool errors
  • Export timeline data for cross-run analysis while preserving links to source evidence
  • Test evaluator agreement by giving independent reviewers the same transcript, vocabulary and configuration

KEEP A HAND ON THE WHEEL

Watch for independent replications across agent systems and domains, measured reviewer time savings, inter-reviewer agreement, false reassurance from shared judge errors, vocabulary sensitivity, performance on much larger transcript collections, secure handling of sensitive logs, offline report behavior, stable provenance links, and evidence that a polished timeline improves decisions rather than merely making them look tidy.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 7, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US