Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
Organizations: UK AI Security Institute
Abstract
Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Input | Output |
|---|---|---|
| Ingest | Logs in a supported format | Scout transcripts representing agent and sub-agent activity recoverable from the source |
| Select | Transcripts and sample/epoch settings | The selected sample and one or more epochs |
| Scan | Selected transcripts, Spec and analysis settings | Results from structural scanners and any enabled judged scanners, saved in a Scout scan store |
| Store | Stored scanner results | Tables of activities, recorded events, token use and audit information |
| Render | Dataframes and source transcripts | HTML reports organised by epoch; tables users can analyse or export |
| Parameter | Default or limit | Control |
|---|---|---|
| Judge models | None; built-in judged scanners disabled | Top level |
| Roll count | 1 | Top level |
| Verifier | On for solo and k-roll; unavailable for cohort | Top level |
| Verifier model | First configured judge model | Top level |
| Verifier spot-check share | 5%; default phase-sampling floor: 3 phases | Top level |
| Phase digests per classification chunk | 40 | Scanner |
| Table | One row represents | Linking fields within a transcript |
|---|---|---|
| transcript_info | Transcript identity and setup | No additional field |
| token_timeline , phase_turns | A model turn: token records or phase attribution | turn |
| flushes , interventions | A recorded or inferred compaction, or human-input event | turn ; several events can share a turn |
| lane_activity | Sub-agent tool activity at a turn | turn , agent_span_id |
| phases | A behavioural phase | phase_index ; turn_start , turn_end locate its range |
| turn_groups | A group of turns within a phase, with narration or fallback text | phase_index , group_index |