WikiSkill keeps a persistent wiki beside agent skills as Apple auto-builds MCP tests
A 29 August arXiv preprint introduces WikiSkill, separating execution traces, accumulated knowledge and executable skills in a persistent wiki that later updates build on. Apple researchers on 28 August described Agent Seer, which generates graded agent tests from Model Context Protocol specifications and covered every tool on 7 MCP specs. Both address how agents retain and verify capability.
Artificial Intelligence··Night
WikiSkill consolidates traces into a persistent agent wiki
A preprint listed on arXiv on 29 August introduces WikiSkill, which keeps a persistent knowledge base beside an agent's reusable skills. Raw execution traces, accumulated knowledge and executable skills are held apart; experience is continuously consolidated into a wiki, and later skill updates build on it. Across the benchmarks and models they tried, the six authors report that this arrangement outscores existing skill-evolution methods. Their ablation puts the weight on accumulation itself: taking the persistent wiki out of the loop weakens skill evolution. They also write that skills transfer across models and model families, and that the skills a model evolves for itself can fall behind those evolved by another model. The work is a preprint that has not been peer reviewed.[1]
Agent Seer builds graded tests from MCP specifications
Apple researchers on 28 August described Agent Seer, which builds test scenarios for tool-using agents from Model Context Protocol specifications instead of hand-written cases. The system enriches function names, descriptions and parameter schemas, produces graded scenarios with synthetic outputs and turns them into multi-turn dialogues. Tested on 7 separate MCP specifications, the pipeline covered every tool on small and medium specifications. The team reports that parameter schema complexity drives quality more than tool-suite size, which comes second. In imperfect scenarios the main failure mode is picking the wrong argument values, a distinction coarse-grained metrics usually miss.[2]
Persistent memory and auto-generated tests target agent reliability
WikiSkill stores evolving skills beside a consolidated wiki the authors say outscores prior skill-evolution methods on their benchmarks. Agent Seer generates specification-driven tests that covered every tool on 7 MCP definitions. WikiSkill comparisons are the authors' own measurements on a preprint; Agent Seer results are Apple's descriptions without independent audit.[1], [2]