AI agents can improve their behavior by developing reusable skills from successful and failed task attempts. But the systems that evolve those skills can lose what they learned along the way, forcing them to rediscover old failures.
WikiSkill, a new framework from Google Research and Virginia Tech, addresses this by adding an organized knowledge layer between an agent’s raw experience and the skills it uses. Instead of repeatedly deriving skills from isolated trajectories, the framework creates a structured wiki from the information it gathers from the agent’s past experiences. It then uses the wiki to build future skills.
In the researchers’ experiments across different domains and models, WikiSkill outperformed existing skill-evolution methods. The results also suggest that its gains can grow with model size and that evolved skills can transfer across models.
For enterprise AI teams, WikiSkill offers a way to turn the execution traces their agents already generate into reusable knowledge and skills.
The missing knowledge layer in skill evolution
Skills package domain-specific instructions, scripts, and workflows into reusable modules. They give agents specialized procedural knowledge without requiring changes to model parameters.
However, creating skills is challenging. Many are still written manually, which means developers must anticipate the workflows and edge cases an agent will encounter. Recent skill-evolution frameworks try to automate that work by running agents on training tasks, analyzing successful and failed trajectories, and generating skill updates from the resulting experience.
Different systems preserve parts of that process. Trace2Skill separately analyzes successful and failed trajectories and merges their lessons into skill patches. EvoSkill searches over candidate skill programs and gives its proposer failure traces alongside a flat history of previous proposals. SkillOpt uses a multi-stage reflection pipeline to analyze trajectories and update a skill document.
Liyan Tang, research scientist at Google and co-author of the paper, says useful diagnoses can disappear even when an optimizer correctly identifies what went wrong. “In many self-improving frameworks, an optimizer might read traces, propose a patch, and then discard the diagnosis, including which fixes failed validation,” Tang said. “That means the system keeps rediscovering the same failures and re-proposing rejected fixes.”
WikiSkill gives those diagnoses and failed interventions a persistent home. Even when a proposed skill is rejected, the system keeps the knowledge that led to it and the result of the validation test, allowing later iterations to build on that history instead of starting from the same information again.
How WikiSkill works
WikiSkill is inspired by Andrej Karpathy’s “LLM Wiki” idea, which proposes compiling experience into persistent, compounding knowledge. Tang describes WikiSkill’s adaptation of the concept this way: “We kept Karpathy’s shape: immutable sources, an LLM-maintained wiki, an index and log, but the input is the agent’s own execution trajectory, and the output is an executable SKILL.md.”
WikiSkill divides the agent workspace into three layers:
The Raw Layer stores immutable execution traces. These include the agent’s actions, reasoning, tool calls, tool outputs, and final answers. It serves as the historical record of what the agent actually did.
The Wiki Layer converts those traces into structured knowledge. It contains individual pages for recurring failure modes and successful strategies, an evolution log, and a skill-impact tracker that records which skill changes were proposed and whether they improved performance.
The Skill Layer contains the procedural instructions available to the agent during task execution. Each skill also links back to the wiki patterns that motivated its creation or modification.

Each evolution cycle has four steps. First, an Inference Agent runs training tasks with the current skills and generates new traces. A Wiki Maintainer then analyzes sampled successful and failed trajectories and updates the wiki. A Skill Proposer reads the updated wiki and selected traces and proposes a new skill or edits to existing ones. Finally, the candidate skill set is evaluated on a validation set. The change is kept only if it improves the best validation score so far.
A case study from ALFWorld, a text-based simulated environment where agents complete multi-step object-manipulation tasks, shows how that works. In the first iteration, the system identified a recurring behavior: the agent repeatedly picked up, examined and returned objects to their original locations. The Skill Proposer created a broad “goal-directed-action” skill, but the skill got rejected during the validation phase.
The wiki retained both the observed behavior and the rejected proposal. In the next iteration, the proposer created a more concrete “break-repetition-loop” skill with the rule “Never Return an Item to Its Origin Location.” That change improved validation performance and was accepted. When later trajectories exposed another looping pattern, the system refined the same skill with an additional rule rather than starting again from scratch.
WikiSkill in action
The researchers evaluated WikiSkill on five benchmarks covering math reasoning, web search, spreadsheets, long-context document question answering, and interactive household tasks. They tested Qwen, Gemma, and Gemini models and compared the framework with Trace2Skill, EvoSkill, SkillOpt, and agents running without skills.
The researchers found that WikiSkill achieved the highest average score for every model tested. Compared with the strongest competing skill-evolution method for each model, its advantage ranged from 3.3 to 12 percentage points.

Against the no-skills baseline, WikiSkill’s average gain grew with Qwen model size. “We saw the opposite of a ceiling: within the Qwen family, gains grew with scale: +12.3, +17.5 and +23.9 points at 4B, 9B and 27B with their own developed skills,” Tang said. At the same time, skills could compensate for some of the gap in model size: Qwen-3.5-9B with WikiSkill reached 47.4% average accuracy, compared with 39.4% for Qwen-3.6-27B without skills.
Skills were also able to transfer across models in some cases. Qwen-3.5-9B scored 70.2% on ALFWorld when given a skill evolved by Qwen-3.6-27B, compared with 63.4% using its own evolved skill.
What enterprise teams can borrow from WikiSkill
The paper provides the evolution algorithm and the prompts used for the Wiki Maintainer and Skill Proposer, making the architecture straightforward to reproduce at a conceptual level.
The main pattern is to preserve full execution traces, extract recurring success and failure patterns into a separate knowledge store, maintain an audit trail of attempted improvements, and use another agent to translate that accumulated knowledge into narrow procedural changes. Those changes should pass an independent validation step before entering the active skill set.
For production systems, the separation between the wiki and the executable skill is also an inference-cost decision. “Memory wants to be exhaustive and a production prompt wants to be lean, so collapsing them requires a less effective compromise,” Tang said. “Instead, we keep the wiki out of the inference agent’s context entirely, so production pays only for compact skills, which is around 45 to 129 lines in our runs.”
The paper’s ablation experiments support that design. In tests using Gemini-3.5-Flash across four benchmarks, the default setup — with the persistent wiki available to the Skill Proposer but not the Inference Agent — averaged 63.7%, compared with 48.7% when the Skill Proposer had no wiki access and the Wiki Maintainer was removed. Giving the Inference Agent access as well reduced it to 60.9%. The researchers suggest that if the inference agent can solve training tasks directly from wiki knowledge, its trajectories reveal less about weaknesses in the skills themselves.
The tradeoff is extra work during the skill-evolution phase. WikiSkill uses a ReAct Skill Proposer, where the model alternates between reasoning and tool calls as it inspects the wiki and execution traces before proposing a change. In the experiments, this took roughly 10 to 20 ReAct turns per evolution iteration, in addition to the Wiki Maintainer call. However, because the researchers processed the full training set in one batch, the number of optimizer LLM calls per iteration did not grow with the number of training examples.
The architecture is most relevant to agents that repeatedly perform multi-step workflows and accumulate enough history for recurring failure modes and successful strategies to emerge. Tang notes that keeping the wiki out of inference can save compute at scale, while more complex applications could eventually benefit from giving the runtime agent access to both skills and selected wiki knowledge. That would require careful decisions about what belongs in each layer so that the wiki complements rather than duplicates the skill files.
There are still production questions the paper does not answer. WikiSkill injects active skills directly into the model prompt, so it does not test skill retrieval or triggering as skill libraries grow. Its validation gate rejects changes that do not produce an immediate performance improvement, even if they could enable later gains. The wiki also grows continuously without an automated pruning mechanism, and the experiments do not cover tasks that run for hundreds of actions or several hours.
The paper identifies automated wiki pruning and online skill adaptation during long-running tasks as open problems. Tang says the team is already exploring that direction: “That transition is something we’re researching right now, actually, and could just be a matter of time before we reach it.”
