Request a scoping call Contact
← Research

Trace an agent evaluation back to the run

The UK AI Security Institute’s open-source Transect makes long agent evaluations easier to inspect. Its source-linked labels can help reviewers find episodes worth checking, but do not validate the labels or the final score.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

An agent can pass a long evaluation and still leave its owners unable to explain how it reached the result. A final score does not show whether the run used a sound sequence of actions, spent much of its budget on setup and rewriting, or encountered a failure the scoring rubric did not capture. For a run that produces hundreds of pages of tool use and sub-agent activity, reading every turn is not a practical review method.

On 7 October 2026, the UK AI Security Institute (AISI) announced Transect, an open-source tool for organising long agent-evaluation transcripts. The accompanying paper was submitted to arXiv on 6 October. Transect places recorded events, token measures, sub-agent activity and model-generated behaviour labels on a shared, source-linked timeline. It offers an evaluation team a way to find episodes worth examining. It does not make the final score more valid by itself.

Put recorded events and judgements on one timeline

Transect starts with a completed run and an evaluation-family specification, or “Spec”, that supplies task context and any activity or sub-agent categories. Its structural scanners extract information already present in the record, such as token use, operator messages, context compactions and sub-agent launches. These run without model calls. Judged scanners ask configured language models to classify activity against the specified vocabulary. Reviewers can move from a label or event back to the source turns, then export the underlying tables for further analysis. The paper describes the pipeline. The implementation README documents supported Inspect logs and OpenClaw telemetry exports.

This division matters. A recorded tool call is evidence that a call appears in the transcript, not proof that its result was useful. A model-generated label is an interpretation of captured material, not an observation. For example, Transect’s sub-agent classifier uses the delegation instructions to describe the work a sub-agent was asked to do. The instruction does not establish what the sub-agent actually did. The common axis follows recorded output turns, not elapsed time. A turn-based chart should not be read as a timeline of real-world duration.

Transect’s report is a view over captured interactions and tool activity, not a special route to hidden reasoning. The distinction helps keep the system boundary honest: what a reviewer can inspect depends on what the harness recorded, the scanners that ran and the transcript material the team is authorised to use.

One large run shows both the value and the limit

The paper demonstrates Transect on one AI research-and-development evaluation that generated almost 13 million tokens. The run included 554 orchestrator tool calls, 79 sub-agent deployments, 57 operator messages and eight context compactions. On its shared timeline, the authors found that only two of 580 orchestrator outputs were labelled “Hypothesising”. That pattern could indicate little hypothesis formation. It could also mean that a scanner assigning one principal activity to each turn misses diffuse or secondary work. The paper presents the finding as a prompt for review, not a settled conclusion about the agent.

The labels were produced in five repeated passes by the same judge model. Across those passes, 474 of the 580 outputs received unanimous research-activity labels. That is a repeatability measure, not proof that the labels were correct. Agreement can show where a model tends to give the same answer. Validity needs a suitable external reference, such as independently collected expert annotations. (Transect paper, Sections 3.1 and 3.5)

The authors say this single-case analysis does not establish the agent’s research capability, the validity of the labels, or causal effects of operator interventions, context compactions or the evaluation scaffold. It also does not measure review time saved, findings missed or improved decisions. Those claims would need multiple samples and epochs, independent reference labels and an explicit statistical analysis. The case demonstrates that a long trace can be organised and followed back to source turns. It does not show that this makes a review faster or more accurate in general.

Keep the trace and its analysis reproducible

The report’s usefulness depends on preserving more than a screenshot or summary. The paper recommends retaining the source transcripts, selected samples and epochs, the family Spec, judging settings, custom analysis layers, stored scan results and relevant software and model versions. These records let a later reviewer distinguish a changed agent run from a changed rubric or analysis pipeline. A fresh output directory alone does not ensure fresh judgements because matching responses may be read from a cache. If repeated calls are intended to measure judge variation, the cache policy must be controlled and verified. (Transect paper, Appendix A.7)

Data handling belongs in the design as well. A judged pass sends the selected transcript material to the configured judge provider. Teams should determine whether those excerpts may leave the evaluation boundary and what retention terms apply. Transect’s exported frames can include task prompts, scaffold instructions, operator messages, delegation text and model-generated explanations. Exporting a dataframe does not make that content anonymous or safe to share. (Transect repository README)

hypothetical fixture

Test whether the report supports a traceable review

Can a source-linked report help reviewers inspect a long agent run without validating its score for them?

Fixture
One representative completed evaluation trace, its recorded events and task brief, with the existing review process available for comparison.
Procedure
  1. Run structural scanners first and check each extracted event against the source log.
  2. Define task-specific activity labels with the evaluation owner and relevant subject experts.
  3. Configure a judged pass, retaining the judge model, prompts, settings and scan results.
  4. Trace sampled labels back to source turns and compare them with independent human review.
  5. Repeat the analysis with a documented cache policy, verify which judge calls are new, then compare reviewer effort and findings across multiple runs.
Result to check
Record which episodes reviewers found, which labels needed correction, and the review time and judge cost. Treat the report as a review queue, not a validated score.
Evidence limit
A single run or agreement between model calls cannot establish label accuracy, causal effects, reviewer benefit or general capability.
Next test
Use several samples and epochs, independent expert reference labels, fixed harness records and a checked cache policy before making comparative claims.

A proposed validation plan for source-linked trace analysis, not a measured result from Transect or an enterprise evaluation.

A hypothetical evaluation checks whether a report helps reviewers locate and inspect important episodes in a long run. It measures review quality and effort against the current method, while treating model labels as candidates for inspection.

Reviewed 2026-10-08

Trial it as a review aid, not as a release gate

For a hypothetical claims-handling agent, a passing average might conceal repeated searches that never obtain the missing evidence, or a handoff that omits a material conflict. An evaluation owner could select a representative run and ask whether a source-linked report helps reviewers find those episodes. Begin with the structural view so the team can see what was recorded before it decides which behaviours to classify. Define a small vocabulary from the task and review criteria, then treat model labels as a way to prioritise transcript excerpts for human inspection.

Compare the report with the existing review method across several runs. Record whether reviewers find the pre-identified issues, which labels they correct, how long review takes and the additional cost of model-judged analysis. Keep the transcript, harness setup, Spec, scanner versions and judge settings with the results. Where the judgement is intended to support a claim about capability or safety, compare it with an independent expert-reviewed sample rather than the judge’s agreement with itself.

The practical decision is whether to add a structured review layer to an evaluation process that has become too large to inspect reliably. The evaluation owner and system owner should first confirm that the trace is complete enough for the question, that reviewers can reach the source records, and that any judged pass fits the data boundary. Expand only if a bounded comparison shows that the report helps reviewers find and explain important episodes. A well-organised trace can make scrutiny more tractable. It cannot substitute for evidence that the agent performed well.

Filed under · Method · Agent evaluation · Observability · Evaluation methods Inference Institute · 08 Oct 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.