From one query to the whole search workflow

A search agent has more to do than compose a good query. Weak initial results call for a revised query, a choice between lexical and vector search, and a decision about when enough information has been found. AWS’s SageMaker AI example puts these decisions into a shared training loop. Its Qwen3.6-27B agent receives a reward for ranking quality across the top ten results at the end of the complete search trajectory. Exceeding the turn or token budget earns a minus-one reward. The developer’s reward design therefore influences both what the agent finds and where its search ends.[1]

AWS’s BrowseCompPlus table makes that stopping decision particularly visible. In the company’s measurement, budget-related failures fall from 22.89% to 0.68%, while nDCG@10 rises from 0.5136 to 0.6354. Average turns decrease from 7 to 6.3. Because a failed run receives a zero ranking score, completing more searches can raise the average even if the documents found do not change. Reaching a usable result has practical value; a maintainer also needs to understand which decision produced that improvement.[1]

I think the useful opening for developers is the ability to tune search quality and execution discipline in the same agent. At least two explanations are plausible, though. Training may teach better queries and tool choices; the budget penalty may also teach the agent to finish otherwise useful searches before they overrun. The fit between training data and each task can change both effects. Tracking ranking quality on successful runs separately from budget-related failures makes the behavior behind an aggregate improvement easier to understand.[1]

A stopping boundary that varies by task

The other tasks show that one stopping policy does not produce the same trade-off everywhere. AWS’s WixQA score increases from 0.5725 to 0.6781, but average turns rise from 4.3 to 4.5. Wands also improves while turns increase from 2.2 to 2.9. FreshStack uses fewer turns but slips slightly from 0.4112 to 0.4089. The balance between a shorter run and a better result varies by task. Squeezing every search into the same turn limit can cut off the value of additional retrieval on some workloads.[1]

Maintenance shifts here from query wording toward the evaluation layer. Developers define tools and rewards, while SageMaker manages training jobs, concurrent rollouts and resumption after interruptions; MLflow provides access to execution traces. Tool choice, revised queries and stopping decisions need to be read together in those traces. Although the base model keeps the same name, training changes its behavior, so assigning the gain solely to a routing layer would be misleading. Discovering whether the reward penalizes costly but useful searches becomes part of tuning the system.[1]

For a small team taking over this agent, I would keep the first experiment narrow: the same questions from its own work, the same tools and turn budget, and versions before and after training. Alongside ranking quality, measure completed runs, tokens spent, latency and results a person has to correct. AWS’s tables do not measure that last review burden. The aim should be to find which tasks justify the cost of additional search rather than select one winning score. Teaching an agent where to stop also gives developers responsibility for setting that boundary around their own work.[1]