Generating and selecting
When a research agent produces many solutions over a day, choosing which one to use becomes part of the experiment. The Malena preprint makes that distinction explicit: it separates the solution selected using validation scores from the best solution researchers can identify after seeing hidden test results. The latter measures the best candidate discovered by the search. The former also tests whether the agent can select that candidate using information available to it.[1]
The highest retrospectively observed score therefore does not represent usable research performance by itself. A wider search may produce a strong candidate without carrying the gain into the final solution if validation fails to select it. Malena’s comparison allows generation and selection to be examined separately. An informative measure of performance on a scientific research task should show both the ceiling among candidate solutions and the information used to make the selection decision.[1]
The authors place this question in comparable conditions across 30 MLE-bench and 40 NatureBench tasks. Main experiments hold the base model, hardware and 24-hour time limit fixed. Malena uses one session with file reading, code writing and command execution; other configurations add search and orchestration components. This comparison helps separate access to tools in the research environment from the contribution of more elaborate orchestration.[1]
The limits of the comparison
Failure to detect a significant advantage for additional layers with strong models does not establish that all configurations are equivalent. The authors report 95 percent confidence intervals, and wide intervals in some comparisons leave uncertainty about the direction and size of differences. MLEvolve’s higher mean estimate with the smaller Gemma 4 model also shows that the choice of model can change the comparison. Reading only the ranking would hide this uncertainty.[1]
The benefit of additional orchestration may diminish as the base model improves. However, differences in tuning effort for external systems may also explain the results, a limitation acknowledged by the paper. Contamination from public competitions is a separate concern. The authors repaired missing test inputs or answer leakage in four tasks, but that intervention does not establish that the model never encountered related tasks during its earlier training.[1]
The methodological question Malena opens is what the selected result represents, before asking how many points additional components produce. Reporting fixed budgets, held-out tests, selection criteria and uncertainty together makes the conditions of an agent’s usefulness easier to assess. NatureBench provides another task set, but success on these two benchmarks does not establish validated benefit across all scientific research. Specifying the scope of the measure preserves the value of the finding.[1]