A route back to history

When a coding agent returns to an old file, a summary of recent turns may leave out the message block it needs. ReCAP makes an interesting choice here: a message block omitted from working context is retained in the archive. The September 30 paper uses filenames and function identifiers in a new request to bring previously omitted messages back. For builders, the change is a recoverable selection of history rather than simply a longer summary.[1]

The mechanism has two layers. Attention computed during ordinary execution supplies historical importance and dependency links. When a request arrives, selection combines those scores with identifier overlap. Selecting a diagnosis can also retrieve the earlier tool output supporting it. Tool calls stay paired with their results, while user instructions and the latest file edits receive explicit protection. This gives a smaller working context a way to preserve relationships between message blocks.[1]

I think reversibility is the useful architectural gain. A message block that appears unimportant today can matter when a later request returns to its file. In one paper example, an initial tool message block remains outside context for seven turns and returns when the user names its functions. Keeping only recent message blocks cannot make that return. Identifier matching may nevertheless favor tasks that explicitly name a file; renamed files or indirect requests may not produce the same selection.[1]

Who can build this memory layer?

That flexibility introduces an integration dependency: the serving system needs access to model attention. ReCAP makes no extra model call when selecting history, but its graph is populated from internal execution statistics. A team receiving only response text from a closed service cannot attach the same mechanism unchanged. Teams serving open models gain an inspectable memory layer while taking on the engineering work of exposing and maintaining those statistics.[1]

Experiments cover two coding benchmarks with Qwen3-Coder and gpt-oss. The authors’ approximately 95 per cent reduction concerns estimated compaction plus cold-restoration latency. It is not a reduction in the duration of the entire development task. Each memory policy also changes subsequent responses and produces its own history, so per-turn token counts are not forced to match. The result supports context selection in these coding conditions without establishing a corresponding reduction in human review.[1]

For a builder adopting the layer, the practical question is whether an old message block can return when needed. Carrying full history, retaining only recent message blocks and selecting with ReCAP are three operating choices. The paper makes the third possible without surrendering the archive. An inspectable graph can explain why a message re-entered context, although an attention score alone does not guarantee that the message block is correct. Designing access to memory, alongside its size, is the useful idea here.[1]