The work behind the manuscript

A research agent’s finished manuscript reveals little about which questions occupied its work. Transect’s useful contribution begins there: it turns a long run into an activity timeline that reviewers can trace back to the original transcript. In the examined run, an agent had to develop a method for detecting distribution changes that harm tabular-model accuracy without seeing ground-truth labels. Seeing the sequence alongside the final submission gives researchers a way to investigate how experiments shaped the proposed method.[1]

The sample consists of one research run. Its 580 orchestrator outputs are classified by research activity and concrete task phase. Structural information appears beside model-generated labels: a subagent launch comes directly from the transcript, while classifying an output as hypothesis formation depends on the judge’s interpretation. The shared horizontal axis follows output order, so a visually long phase cannot be translated into an equally long duration. That distinction sets an initial boundary on what the display measures.[1]

Two explanations for a missing label

Only two outputs receive the hypothesis-formation label, while none is labelled analysis design. A first reading suggests an agent occupied with operations and writing rather than searching for methods. But the scanner assigns one principal activity to each output. Ideas distributed through an implementation discussion could disappear under that rule. The same observation could arise from a narrow research process or a narrow classification scheme. The reviewer’s question becomes concrete: which new hypotheses were actually proposed and tested in those transcript passages?[1]

Token accounting adds another view. Of roughly 14.5 million tokens attributed to new work, 7.8 million go to self-review and 2.9 million to manuscript preparation. Literature activity also concentrates largely in the writing phase. This distribution invites separate judgments about effort spent on the research question and effort spent refining the text. Token counts still measure thought quality indirectly: a short methodological idea could matter more than a long writing process. The spending pattern helps an expert locate material relevant to judging the eventual scientific contribution.[1]

Seeing the same result again

Research-activity classification is repeated five times with the same model; labels agree on 474 outputs. That stability helps another evaluator reconstruct the analysis. Validity concerns whether experts would assign the same categories to those passages. One judge model could consistently repeat a shared error. Judgments from different models and independent expert labels are among the routes described for examining those questions separately.[1]

For me, Transect’s strongest result is a better case study before a larger verdict about research capability. In the next comparison, researchers could measure writing and idea-formation shares across multiple runs in the same task family. Checking the labels against independent expert judgments can help distinguish a recurring emphasis on writing from a feature of this setup. Gaps in one run can then be investigated without immediately converting them into a capability score. The person assessing the finished manuscript gains a clearer map of which work steps support which scientific claims.[1]