Eigen RadarAI
Analysis

Only 62-87 per cent of captured flags came from an exploit the trace can show

CTF-ABACUS, published on arXiv on 29 August, reconstructs paths to flags in capture-the-flag challenges step by step; across 1435 attempts trace-verified exploits made up only 62-87 per cent of recovered flags. Apple's Agent Seer description on 28 August builds tests from Model Context Protocol specifications and reports covering every tool on seven separate specs. One addresses score inflation, the other trustworthy testing against the same agent measurement problem.

Artificial Intelligence··Evening
In a bright laboratory, a blue resin block with a large backlit hollow interior sits on a precision balance; three differently shaped resin blocks rest on the bench.

CTF-ABACUS separates flag paths step by step

CTF-ABACUS, published on arXiv on 29 August, reconstructs how language-model agents reach the flag in capture-the-flag challenges, step by step. The framework separates a genuine exploit from direct flag exposure, memorised recall, external lookup, guessing and unsupported claims. The team examined 1435 attempts across six models and 240 challenges, producing 2870 solve profiles. Trace-verified exploits make up only 62-87 per cent of recovered flags across the benchmarks, and shortcut recoveries follow substantially shallower trajectories. The authors conclude that current benchmarking overstates agent offensive capability. The work is a preprint without peer review, and the profile classification rests on rules the authors define rather than an independent audit or peer-reviewed validation, with no external replication, outside audit, independent confirmation, third-party review, independent audit, external validation, field confirmation or independent review reported for the headline percentages or shortcut trajectories.[1]

Agent Seer builds tests from MCP specifications

Apple researchers on 28 August described Agent Seer, which builds test scenarios for tool-using agents from Model Context Protocol specifications instead of hand-written cases. The system enriches function names, descriptions and parameter schemas, produces graded scenarios with synthetic outputs and turns them into multi-turn dialogues. Tested on 7 separate MCP specifications, the pipeline covered every tool on small and medium specifications. The team reports that parameter schema complexity drives quality more than tool-suite size, which comes second. Comprehensive tool descriptions consistently reduced failures, while few-shot prompting was followed by severe inaction in some models. The measurement ran against a server mimicking a proprietary tool rather than the tool itself, and the authors say the main failure mode in imperfect scenarios is picking wrong argument values rather than refusing to act at all, with no independent audit, peer review, outside confirmation or third-party review of the generated tests reported.[2]

Score inflation and automated tests answer the same measurement problem

CTF-ABACUS shows only 62-87 per cent of flag scores came from genuine exploits while Agent Seer aims to generate comprehensive tests from MCP specs. One separates why current scores inflate, the other tries to automate trustworthy testing. CTF-ABACUS worked across 1435 attempts and 2870 solve profiles on six models and 240 challenges; Agent Seer covered every tool on 7 MCP specifications. CTF-ABACUS results rest on the authors' rules; Agent Seer was measured on a mimicking server on 28 August. Shortcut recoveries follow shallower trajectories; on Agent Seer comprehensive tool descriptions consistently reduced failures. For readers the concrete development is agent capability being re-weighed in both security-challenge scores and tool-calling exams in the same week, with one line exposing shortcut inflation and the other automating specification-driven tests without independent replication, peer review, outside validation, independent audit or third-party confirmation on either side.[1], [2]

References

  1. News sourcearXivOnly 62 to 87 percent of captured flags came from an exploit the trace can show↩1↩2
  2. News sourceApple Machine Learning ResearchApple describes a pipeline that builds agent test scenarios from MCP specifications↩1↩2