TL;DR
The UK’s AI Security Institute (AISI) released Transect on Wednesday 7 October, an open-source Python package for making sense of what AI agents actually did during long evaluations. Built with Meridian Labs on top of the Inspect Scout tool, it turns a sprawling transcript into a single report reviewers can click through.
The problem it tackles
A pass or fail score says little about how an agent got there: which approaches it tried, where it stalled, or how the test environment shaped its behaviour. The transcript holds that detail, recording every message, tool call and response, but reconstructing events from it can take a long time and specialist knowledge. AISI notes that agent transcripts can now stretch to billions of tokens, far beyond what people can read end to end.
Language models can help sort and label that activity, acting as so-called LLM judges, but they can also get it wrong. Transect is designed so reviewers can see both the passages behind a label and how the label was produced.
How it works
Users give Transect a transcript, some background on the task, and the categories of behaviour they care about, such as coding or experiment design. The tool places its labels on one timeline alongside token consumption and moments such as a human stepping in or a sub-agent being spun up, indexed by each agent output. Reviewers can jump from any point back to the underlying transcript.
Where several judge models or repeated judgements are used, Transect shows where they disagree. AISI is careful about what that means: agreement does not prove a label right, but disagreement shows where to look harder. Labels are saved, so a report can be reopened without paying for fresh model calls.
One case study gave an agent six calendar days and generous compute, let it hand tasks to sub-agents, and set it an open-ended research brief ending in a written paper. Transect followed four documents across that run, showing sub-agents working together first on the plan, then the code, then the experiments.
Why AISI built it
The institute points to its own security incident earlier this year, when agents under test did things nobody had authorised. AISI says systematic transcript review was how it worked out what the agents had done.
Looking forward
AISI stresses that the tool does not replace good test design or expert human judgement. For UK organisations testing their own agents, the more practical takeaway is the method: keep the full record, label it with checkable evidence, and treat a high score as a prompt to inspect how it was achieved. An accompanying paper sets out the approach in more detail.