Eigen RadarAI
Analysis

WikiSkill keeps a persistent wiki beside agent skills as Apple auto-builds MCP tests

A 29 August arXiv preprint introduces WikiSkill, separating execution traces, accumulated knowledge and executable skills in a persistent wiki that later updates build on. Apple researchers on 28 August described Agent Seer, which generates graded agent tests from Model Context Protocol specifications and covered every tool on 7 MCP specs. Both address how agents retain and verify capability.

Artificial Intelligence··Night
In a warmly lit workshop lined with worn wooden card drawers, a matte robot arm lifts one blank card from an open drawer into the light while its second tool rests by the bench.

WikiSkill consolidates traces into a persistent agent wiki

A preprint listed on arXiv on 29 August introduces WikiSkill, which keeps a persistent knowledge base beside an agent's reusable skills. Raw execution traces, accumulated knowledge and executable skills are held apart; experience is continuously consolidated into a wiki, and later skill updates build on it. Across the benchmarks and models they tried, the six authors report that this arrangement outscores existing skill-evolution methods. Their ablation puts the weight on accumulation itself: taking the persistent wiki out of the loop weakens skill evolution. They also write that skills transfer across models and model families, and that the skills a model evolves for itself can fall behind those evolved by another model. The work is a preprint that has not been peer reviewed.[1]

Agent Seer builds graded tests from MCP specifications

Apple researchers on 28 August described Agent Seer, which builds test scenarios for tool-using agents from Model Context Protocol specifications instead of hand-written cases. The system enriches function names, descriptions and parameter schemas, produces graded scenarios with synthetic outputs and turns them into multi-turn dialogues. Tested on 7 separate MCP specifications, the pipeline covered every tool on small and medium specifications. The team reports that parameter schema complexity drives quality more than tool-suite size, which comes second. In imperfect scenarios the main failure mode is picking the wrong argument values, a distinction coarse-grained metrics usually miss.[2]

Persistent memory and auto-generated tests target agent reliability

WikiSkill stores evolving skills beside a consolidated wiki the authors say outscores prior skill-evolution methods on their benchmarks. Agent Seer generates specification-driven tests that covered every tool on 7 MCP definitions. WikiSkill comparisons are the authors' own measurements on a preprint; Agent Seer results are Apple's descriptions without independent audit.[1], [2]

References

  1. News sourcearXivSmaller models with an accumulated wiki outscore larger models without one↩1↩2
  2. News sourceApple Machine Learning ResearchApple describes a pipeline that builds agent test scenarios from MCP specifications↩1↩2