Transect maps research-agent activity onto a traceable timeline
Transect is an open-source tool that places tool use, task phases and token consumption from long AI research runs on one timeline. Its single-run preprint examines 580 orchestrator outputs. The new-work accounting totals approximately 14.5 million tokens. Self-review takes 7.8 million of them, and manuscript preparation takes 2.9 million. Labels link back to the original transcript turns for detailed inspection.
Artificial Intelligence··Evening
Labels link back to original transcript turns
Transect is an open-source tool for examining long AI-agent runs. The single preprint, by Cozmin Ududec, Alexandra Abbas, Konstantinos Voudouris and Toby D. Pilditch, brings tool activity, token use and behavioral labels onto one timeline. Built on Inspect Scout, the package links each label to its original transcript turn. Users define task context and behavioral categories, then reuse the configuration across different runs.[1]
The demonstration classified 580 outputs five times
The demonstration covers one run from the CRUX research initiative. The agent’s task is to develop a method that detects harmful changes in the data distribution affecting tabular models, without requiring labels. It must also prepare a package for reproducing the work. A human operator helps with operational problems. Claude Opus 4.7 classifies the run’s 580 orchestrator outputs five times. Research-activity labels received unanimous judgments on 474 outputs. Two outputs were labelled hypothesis formation. The authors caution that assigning a principal category to each output can leave secondary activities and dispersed work outside the classification.[1]
Self-review and writing dominate token use
The new-work accounting totals approximately 14.5 million tokens. The allocation to self-review is 7.8 million, with 2.9 million used for preparing the manuscript. Tokens are text units processed by a model. This accounting combines inputs and outputs with cache writes capped by context growth. Structural scanners read tool activity and operator messages without calling a model, while model-based scanners assign task-phase labels. The horizontal axis follows the sequence of orchestrator outputs, rather than time elapsed. Inputs, prompts and the judge’s response remain available alongside the labels. Users can export tables for further analysis.[1]
Related columns
For more information on this topic, you can read the related columns.