WikiSkill keeps what agents learn and lifts performance across five benchmarks
Google Research's WikiSkill framework keeps what agents learn in a persistent wiki-style knowledge base and reports double-digit gains across five benchmarks. Earlier methods discarded lessons at the end of each cycle; WikiSkill reuses traces of failures and successes on later attempts. Gemini-3.5-Flash moved from 49.5 per cent to 68.1 per cent and Qwen-3.6-27B from 39.4 per cent to 63.3 per cent in reported measurements, aiming to raise performance on repeated tasks without retraining the models.
Artificial Intelligence··Midday
WikiSkill accumulates lessons where earlier methods discarded them
Google Research's WikiSkill framework lets AI agents keep what they learn across runs instead of discarding it when a task ends. Crypto Briefing reported that earlier skill-evolution methods threw away the lessons drawn during a run at the end of each cycle. WikiSkill instead stores the trace of failures and successes and reuses that material on later attempts. Creati added that the framework aims to raise performance on repeated tasks without retraining the underlying models, and that the agent refers to the mistakes it met and the steps that worked in earlier runs on later tasks.[1], [2]
Raw, wiki and skill layers turn traces into procedures
WikiSkill rests on three layers. Creati described a raw layer that holds execution traces unchanged, a wiki layer that turns that raw material into accumulated knowledge, and a skill layer that holds the procedures the agent can run. Crypto Briefing wrote that the framework gathers what agents learn in a persistent wiki-style knowledge base. The agent accumulates traces of failed attempts and steps that worked in this structure and draws on those layers on later tasks.[2], [1]
Reported measurements show double-digit gains on two models
In the reported measurements Crypto Briefing carried, Gemini-3.5-Flash moved from 49.5 per cent to 68.1 per cent and Qwen-3.6-27B from 39.4 per cent to 63.3 per cent. The same outlet wrote that Google Research reported double-digit gains across five benchmarks. Creati reported that the framework aims to raise performance on repeated tasks without retraining the models and that the agent stores the mistakes from earlier runs and the steps that worked in this structure.[1], [2]