Eigen RadarAI
Analysis

BOTTLED finds most agent-built solutions trail direct model performance

The BOTTLED preprint tests whether language-model agents can build reusable solutions for workloads containing millions of examples. Of 60 runs across three tasks, 48 fell below the lower confidence bound of the corresponding model’s direct responses. Agents could write programs, train small models or create hybrid solutions. The single study compares those approaches under specified time and token budgets.

Artificial Intelligence··Evening
An exposed circuit module in a metal enclosure connects by a short cable to a laptop on an engineering workbench.

Reusable solutions were compared across 60 runs

BOTTLED evaluates whether language-model agents can develop reusable solutions for large workloads. The single Agent in a Bottle preprint tests ten models twice on each of three tasks. The authors report that 48 of the 60 results fell below the lower bound of the 95 per cent confidence interval for the corresponding model’s direct responses. The stronger of two small models trained under the same token budget also outperformed 31 agent runs.[1]

Agents receive millions of examples at the outset

The complete set of unlabelled examples is available to an agent in a file before work begins. There is no prescribed solution format. Training a smaller model, writing a program with rules and building a retrieval index are all allowed. A hybrid can instead reserve large-model calls for difficult cases. Tasks involve extracting attributes from product descriptions, classifying shopping query–product relevance and distinguishing human-written from model-generated text. Attribute extraction involves approximately 4.77 million examples. The relevance workload contains about 2.62 million examples, and text detection has about 5.62 million. For evaluation, agents must submit a file with their predictions for the complete workload.[1]

Time and token limits left some outputs unfinished

Each run receives 10 hours and 5 million weighted tokens. Tokens are units of text processed by a model; outputs, uncached inputs and cached inputs have different accounting weights. An agent can call only its own model’s application interface. Eight of the 21 runs terminated at the budget limit failed to finish the required output. An additional experiment removed that termination but retained the overall performance gap. Successful lower-cost solutions were obtained on some tasks. The code was released under the MIT licence.[1]

References

  1. News sourcearXivBOTTLED tests whether agents can turn knowledge into reusable solutions for millions of tasks↩1↩2↩3