The path of an experiment

Must a coding agent reread its previous conversation when returning to a long optimization task? Stateless Language Agents, or SLA, answers with a different operating arrangement. Its proposed architecture stores candidate code, experimental outcomes and failed attempts in an external harness. Each agent invocation begins in a fresh session. The interesting choice for builders is retaining accumulated work without carrying the entire conversation that produced it. Durable progress no longer depends on continuing that conversation. The harness becomes responsible for giving the next session a usable starting point.[1]

Follow an experiment through that arrangement. The advisor receives the task, the best candidate code and an outcome summary grouped by research direction; it uses these to assign concrete experiments. A worker sees its assignment, retained candidate and recent local feedback. It cannot access other workers' directories or the full history. An evaluator outside the worker environment scores its candidate. The outcome returns to durable state and supplies the starting point for a subsequent invocation. Continuing the work therefore does not require continuing the session that proposed it.[1]

What an inconclusive trial means

For maintainers, the useful distinction is between code that failed to run and an idea that failed to improve the objective. The advisor's summary separates measured non-improvements from trials left inconclusive by implementation, runtime or correctness failures. An untested direction is not an exhausted direction. I think this is where the new responsibility sits: the builder designs not only which messages to summarize, but how to classify why an experiment produced no usable result. A mistaken classification can reproduce the same mistaken decision even in a fresh session.[1]

The researchers test software engineering, kernel optimization and algorithm design, with cumulative budgets reaching one billion tokens. Experiments restarted from shared checkpoints separately remove advisor context reconstruction, worker isolation and explicit assignments. That is a useful comparison for inspecting the architecture's components. Yet task selection or long budgets favoring particular methods may also influence the results. Most configurations receive only one run, limiting confidence when transferring the effect to another project. These are executable optimization tasks, rather than evidence that the arrangement improves an ordinary team's development workflow.[1]

Owning durable work

Across the measured configurations, the advisor consumes 0.24–0.51% of tokens and 1.2–2.3% of model cost. This shows that a separate coordination layer can operate with modest model expenditure in these experiments. Computation used to execute the experiments is excluded from that accounting. Human review, debugging and harness maintenance cannot be inferred from those percentages either. For a small team, an inexpensive advisor does not establish a cheaper overall workflow. The team still needs to locate the expense that has moved into execution, evaluation and maintaining durable state.[1]

The builder gains the option to replace advisor and worker sessions independently of their conversations. In exchange, candidates, outcomes and recovery points need reliable storage. I would begin with one optimization task that already has executable tests. Keeping the model, seed code and evaluator fixed would let a team inspect which candidate returns after a session interruption. The useful opening is the ability to explain where unfinished work resumes, before making any promise about longer autonomous development. SLA places that responsibility in the operating arrangement the builder maintains, giving the team a concrete point of control and a concrete maintenance obligation.[1]