Research series · 2026 · Report 08 of 16
Report 08 of 16 · Part II · Non-parametric · Skills

Skills and tools

the self-extending agent

Skills turned procedural learning into files an agent can write, share, and load on demand, and the 2026 evidence shows real compounding within a domain and transfer across models without any weight update. At library scale the problems changed: relevant skills cause most skill-induced failures, per-skill effects vary by task in ways aggregate scores hide, every repository release invalidates part of the library, and the ecosystem grew faster than its security. The skill is no longer the hard part; the loop that decides which skills exist, load, and survive is.

ContentsReport 08 · Non-parametric
$ tree ./08-skills-tools
./08-skills-tools
├── 01-when-the-right-skill-makes-the-agent…# 352 words · 1 figure
├── 02-from-voyager-to-a-shared-package-format# 529 words
├── 03-how-the-self-improvement-loop-works# 1034 words
├── 04-improving-the-improver# 274 words
├── 05-what-the-evidence-shows# 494 words
├── 06-where-it-breaks-at-library-scale# 951 words · 1 figure
├── 07-state-of-play-2026# 301 words
├── 08-what-ships# 670 words
└── 09-open-problems# 404 words
9 sections · 42 references · 2 figures
31.5%
skill activation after gating, accuracy unchanged
78 of 1,000
replayed chats that picked a skill production blocks
307
failures traced to skills, most of them relevant-looking
42
references · 17 from Jun 24 – Sep 24, 2026
21 min
reading time
Storage
Non-parametric: code functions, SKILL.md folders, MCP servers, tool wrappers
Engines
LLM reflection over traces, failures, and source code; bandit and population search; RL for selection and distillation (hybrid); recursive meta-skills
Evaluator
Execution and deterministic verifiers (best); held-out dev splits; learned activation gates; co-evolved surrogate verifiers; self-feedback (worst)
Loop timescale
Per task (activation, test-time synthesis) to per release (maintenance) to per library generation
Loop closure
Research: mostly closed loop on benchmarks. Practice: every public skill edit studied so far is human-merged; agent-written skills run with the user on the loop
Evidence maturity
Mostly 2026 preprints, a few peer-reviewed (SkillLearnBench, CoEvoSkills); strong within-domain results; failure, security, and maintenance studies concentrated in July–September 2026

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 09

#When the right skill makes the agent worse

In August 2026, Dong et al. ran agents on SkillsBench and SWE-Skills-Bench twice per task: once with a loaded skill, once without it or with a semantically matched reference skill. Wherever the skill-guided run did worse, they traced the regression to the skill. They found 307 skill-induced failures: 125 functional failures, where the task broke, and 182 efficiency regressions, where the task succeeded at a higher cost 1. Irrelevant skills were not the main culprit. Seemingly relevant skills caused most of the functional failures.

The efficiency taxonomy shows how. The largest category, Excessive Verification, accounts for 67 cases; the next, Heavy Implementation Pipeline, accounts for 30 1. In both, the skill was on topic and its advice was reasonable in general. The agent read a validation checklist as a required procedure and ran every check, or read a construction recipe (subprocess workflows, multi-stage conversion, runtime simulation) as the only acceptable path and built all of it for a task that needed a fraction. In the authors' words, skills "often turn validation checklists and construction recipes into mandatory work" 1. An irrelevant skill gets ignored. A relevant skill gets obeyed, including the parts that do not apply.

Fig 08.1 · Where skill-induced harm comes fromDong et al. · Aug 2026
307 SKILL-INDUCED FAILURES · SKILLSBENCH AND SWE-SKILLS-BENCH 125 functional: the task broke 182 efficiency: the task passed at a higher cost Most caused by seemingly relevant skills, not irrelevant ones 67 Excessive Verification 30 Heavy Implementation Pipeline other categories
Relevant skills do most of the damage. Running each task with and without a loaded skill traced 307 regressions to the skill: 125 functional failures and 182 efficiency regressions, where the task still passed 1. Excessive Verification (67) and Heavy Implementation Pipeline (30) are the two largest efficiency categories, and both come from on-topic advice read as mandatory work; an eval that scores only pass rate would record the 182 as neutral or positive 1.

That result inverts the obvious failure model for skill libraries. The obvious worry is retrieval: load the wrong skill and the agent goes astray, so improve retrieval. The 307-failure study says a large share of the damage comes from loading the right skill and following it too literally. Better retrieval makes that worse, because it loads more relevant skills more often. Within six weeks, two papers answered with the same design move: treat "is this skill relevant?" and "should this skill run now?" as separate decisions, and learn the second one from matched with/without evidence 2,3.

Library-scale failures of this kind are a 2026 phenomenon. Before that, the skills literature was about getting the loop to work at all: can an agent write reusable procedures, verify them, and get better by reusing them? It can, and the way that loop was built explains why its failures showed up only once libraries got large, shared, and long-lived.

Section 02 / 09

#From Voyager to a shared package format

A skill, in the sense used here, is procedural knowledge an agent writes to an external artifact that a later run of the same or a different agent can retrieve and execute: a code function, a tool wrapper, an MCP server, or a folder with a SKILL.md file. The learning is non-parametric: the weights never change. What changes is the set of procedures the agent can load. Skills differ from experiential memory (see 07, Experiential memory) in that memory stores what happened and skills store how to do something; they differ from harness edits (see 09, Harness self-improvement) in that a skill changes what the agent knows how to do, not the control flow that runs it.

The idea predates the format by two years. Voyager (2023) wrote JavaScript skills in Minecraft, admitted only code that passed execution and self-verification checks, retrieved skills by embedding, and collected 3.3× more unique items than prior agents 4. LATM (2023) split tool making from tool use: GPT-4 wrote tools that GPT-3.5 called at comparable accuracy and lower cost 5. By 2025, SkillWeaver's web agents distilled practice into API functions that lifted weaker agents by up to 54.3% relative on WebArena 6, and Alita generated MCP servers at runtime to reach 75.15% pass@1 on GAIA validation 7. Two design choices from that lineage still hold: executable skills can be gated by execution, and libraries are retrieved rather than loaded whole. One gap also carried forward: none of these systems deleted, retired, or re-validated a skill.

Until late 2025, every system invented its own skill format, so nothing one agent learned could be used by another agent's harness. Anthropic's Agent Skills, launched on October 16, 2025, fixed the format 8. A skill is a folder with a required SKILL.md file carrying YAML frontmatter (a name and a description) and optional scripts/, references/, and assets/ directories. Loading uses progressive disclosure: the agent reads only names and descriptions at startup, loads the full instructions when a skill activates, and pulls resources on demand. Anthropic published the format as an open standard in December 2025 9, and Codex, Gemini CLI, Cursor, Windsurf, and GitHub Copilot adopted it 10. A research system that emits a SKILL.md now produces something every major coding agent can load, and a registry can index skills without knowing which agent wrote them. Progressive disclosure made large libraries possible in principle, because the per-skill startup cost is a description, not a document. It also moved the selection problem into the description field, which is where much of the later trouble sits.

The standard is thin on purpose. It says where instructions live, not whether they are correct, current, safe, or worth loading. A June 2026 essay, "Skill Libraries Need CI, Not More Prompts", argued that once skills can be installed, auto-selected, and updated, they behave like governed persistent state, and that the commit gate matters more than the proposer 10. The essay covered lifecycle tooling (SkillOps, PACE, SkillRevise, Workflow-to-Skill, MemoRepair, and others). It did not cover the mechanics by which agents write and improve skills for themselves, or the evidence gathered from June through September 2026 about what happens when those libraries get large.

Section 03 / 09

#How the self-improvement loop works

Every self-extending agent runs the same loop: gather evidence, propose a new or revised skill, test it, commit or reject, and retrieve it later. The systems differ in what they mine, how they propose, and what decides.

Mining code before experience: Code2Skill

The newest extraction work skips agent experience altogether. In September 2026, Code2Skill turned code units from 19,769 popular, actively maintained GitHub repositories into skill records (atomic operations, composite workflows, recurring patterns), verified each one by reconstructing it without seeing the source body and then comparing against the source, and kept 1,006,822 accepted records with provenance metadata 11. Across 72 protocol-matched evaluations covering nine model settings and eight benchmarks, retrieving from that bank improved models by 11.7% on average over matched baselines and beat them in 57 cases; under a shared interface it also beat trajectory-derived skill banks on all seven shared benchmarks 11. Repository-derived skills solve the cold start that trace mining cannot, because they exist before the agent has done anything. They also pin a million records to code that will change, which makes staleness (see Where it breaks) more pressing.

Consolidating many traces at once: Trace2Skill

Trace mining still carries the strongest transfer evidence. Early extraction was sequential: finish an episode, write or edit a skill, move on, which overfits to whichever episode came last. Trace2Skill (March 2026, v5 June 2026) consolidates many trajectories in parallel instead, inducing recurring patterns across a broad set of traces, resolving conflicts between them, and writing a unified skill directory; it outperforms both sequential skill editing and ReasoningBank-style retrieval of raw experience 12.

Its most important result concerns transfer. Skills that Qwen3.5-35B evolved on its own trajectories improved a Qwen3.5-122B agent by up to 57.65 absolute percentage points on WikiTableQuestions, with no weight update 12. A smaller model's procedural lessons, written as files, made a larger model much better at the same domain. The comparison is within a model family, "up to" marks a best case, and the paper is labeled work in progress; within those limits, it is the strongest evidence in the series that a lesson written into text survives a change of model. That is the portability property the series keeps returning to (see 12, Consolidation and co-evolution): code and structured procedure transfer across models where tuned prose often does not.

Mining failures: EvoSkill and SkillRevise

Successful traces show one path that worked. Failed traces show where the agent's default behavior breaks, which is what a skill should fix. EvoSkill (Sentient AI and Virginia Tech, March 2026) analyzes execution failures, proposes new skills or edits, and keeps only those on a Pareto frontier of held-out validation performance, with the model frozen throughout 13. On OfficeQA, exact match rose from 60.6% to 67.9%; on SealQA, from 26.6% to 38.7%; and skills learned on SealQA transferred zero-shot to BrowseComp for a 5.3% gain 13. Merging libraries from independent runs beat any single run (67.9% against 64.5%), and gains flattened as the training fraction grew 13. SkillRevise (HKUST, May 2026, v4 September 2026) handles the cold start with one imperfect skill: it diagnoses defects from execution evidence, retrieves general repair principles, applies edits anchored to the observed failure, and keeps the best re-executed candidate; on SkillsBench, base-agent success rose from 36.05% to 61.63% 14.

The optimizer race: SkillOpt and its September challengers

SkillOpt (Microsoft, May 2026) treats skill text as trainable external state 15. It proposes bounded edits, keeps a history of rejected edits so it does not cycle back to them, and commits an edit only when held-out validation improves. Across 6 benchmarks, 7 target models, and 3 harnesses (direct chat, Codex, Claude Code), it was best or tied in all 52 cells; with GPT-5.5 it added 23.5 points in direct chat, 24.8 inside Codex, and 19.1 inside Claude Code 15.

SkillOpt became the baseline to beat within four months. In September 2026, COBRA-Skills cast skill optimization as budgeted sequential search, using a contextual bandit to decide which candidates deserve an expensive evaluation; it reported the strongest average across six benchmarks and three target models at 55–58% lower optimization cost than SkillOpt, using 50 optimization examples per benchmark 16. GraphSkillEvo represented a skill as a graph of steps and transitions and evolved a population with mutation and crossover, beating SkillOpt by 4.01% average accuracy on GPT-5.4-nano and 1.76% on GPT-5.4 across five benchmarks 17. SkillAdam added an optimization memory and a volatility-driven edit budget, functional analogues of Adam's two moments, so that iteration-local feedback stops overwriting earlier corrections 18. The proposer is getting cheaper, more structured, and more stable. None of the three changes the acceptance rule: an edit survives because a validation score went up.

Who judges the skill

A skill loop is bounded by its evaluator, and the August 2026 evidence questions whether explicit skills are doing the work at all. ContinualSkillBench built five domains of 100 interconnected subtasks each, ordered by difficulty with chances for cross-task reuse 19. Sequential execution generally improved performance, but plain in-context learning performed comparably to explicit skill maintenance on average, "suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone"; explicit skills helped selectively on tasks needing reusable procedures or precise outputs, and less capable models accumulated larger, more fragmented collections of task-specific skills 19. The result echoes CL-Bench's finding for memory systems (see 07, Experiential memory, and 14, Measuring self-improvement).

Earlier 2026 benchmarks point the same way. SkillLearnBench (CMU and Amazon AGI; COLM 2026) found that every continual skill-learning method beat the no-skill baseline, but none led consistently, stronger backbones did not reliably produce better skills, and self-feedback alone induced recursive drift, while repeated external feedback produced real improvement 20. A May 2026 lifecycle study found that a model can be a strong skill extractor and a weak skill consumer, and that model-generated skills show non-trivial negative transfer even when they help on average 21. Where no ground-truth tests exist, CoEvoSkills (COLM 2026) evolves a Surrogate Verifier alongside the Skill Generator without showing it test content, and beat five baselines on SkillsBench for both Claude Code and Codex 22. The ordering from the series argument holds: external checks beat co-evolved judges, which beat self-judgment.

Section 04 / 09

#Improving the improver

Every system so far has a fixed improvement procedure written by its authors: how to analyze a failure, which traces to retrieve, how much budget to spend, how to phrase an edit. MetaSkill-Evolve (July 2026) makes that procedure a skill too 23. Its improvement pipeline has five roles (Analyzer, Retriever, Allocator, Proposer, Evolver), and a meta-skill encodes how each should operate. Task skills evolve on a fast loop from execution traces. The meta-skill evolves on a slower loop, "under the same pipeline applied to itself, with no additional model or objective" 23. All five roles run on the same frozen backbone.

Over the raw backbone, MetaSkill-Evolve added 23.54 points on OfficeQA, 16.09 on SealQA, and 1.92 on ALFWorld, and it beat no-skill, static-skill, and single-level evolution baselines on all three 23. Beating single-level evolution is the result that matters: it says evolving the procedure adds something over evolving the skills with a fixed procedure.

The recursion is real and bounded. The meta-skill improves the improvement procedure, and it is improved by that same procedure; there is no third level, and nothing in the paper shows the two-level process converging or continuing to improve beyond the reported runs. The spread across benchmarks (23.54 points on OfficeQA, 1.92 on ALFWorld) suggests the meta-level helps most where the base skill loop has the most room. Two timescales also recur across the series: explore fast in context, consolidate slowly. Here the slow store is still text, a meta-skill rather than weights, so the whole recursive stack stays inspectable and can be rolled back. Report 13 (Recursive self-improvement) covers the definitional debate about whether this counts as recursive self-improvement.

Section 05 / 09

#What the evidence shows

Across the 2026 literature, gains are large when skills target a specific domain and are checked by execution, and smaller and more variable when the skill set is broad or the check is weak.

Date System Setting Baseline → result What it demonstrates
Sep 2026 SkillApt 3 SRA-Bench Same accuracy as BM25 top-1 (0.838 vs 0.838); skill activation 100% → 31.5%; mean tokens −74.3% Deciding whether to load a retrieved skill is a separate, learnable decision
Sep 2026 Code2Skill 11 72 evaluations, 9 model settings, 8 benchmarks +11.7% average over matched baselines; better in 57/72; beats trajectory-derived banks on 7/7 Skills mined from 19,769 repos work before any agent experience
Sep 2026 COBRA-Skills 16 6 benchmarks, 3 target models Strongest average among compared methods at 55–58% lower optimization cost than SkillOpt The proposer side is getting cheaper
Sep 2026 OpenClaw registry study 24 61,990 skills covered by three scanners Scanners disagree on 23,702; weighted sensitivity 21.67%–61.06% Registry security cannot rest on one scanner
Aug 2026 Repo2Skill-Evo 25 57 repos, 105 release transitions Every transition invalidates part of the skills; frontier agents 29.9%–69.7% macro F1 at maintenance Skills decay with every release
Aug 2026 EVOMAL 26 6 models, 153 SWE-bench Verified tasks 20.3%–41.8% of tasks add a new malicious skill; poisoned libraries hold 4.9–9.0× the planted count Self-authoring turns one planted skill into a worm
Aug 2026 Harmful-skills study 1 SkillsBench, SWE-Skills-Bench 307 skill-induced failures (125 functional, 182 efficiency) Relevant skills cause most failures
Aug 2026 ContinualSkillBench 19 5 domains × 100 subtasks In-context learning comparable to explicit skill maintenance on average Much apparent skill learning is context adaptation
Jul 2026 MetaSkill-Evolve 23 OfficeQA / SealQA / ALFWorld +23.54 / +16.09 / +1.92 pp over raw backbone Evolving the improvement procedure adds over single-level evolution
May 2026 SkillOpt 15 52 model/benchmark/harness cells Best or tied in 52/52; GPT-5.5 +23.5 (chat), +24.8 (Codex), +19.1 (Claude Code) Gated edit loop beats other optimizers and human skills
May 2026 SkillRevise 14 SkillsBench 36.05% → 61.63% base-agent success Execution-anchored repair from a cold start
Mar 2026 Trace2Skill 12 WikiTableQuestions Qwen3.5-35B-evolved skills lift Qwen3.5-122B by up to 57.65 pp Skills transfer across model scale without weight updates
Mar 2026 EvoSkill 13 OfficeQA; SealQA → BrowseComp 60.6% → 67.9%; 26.6% → 38.7%; +5.3% zero-shot on BrowseComp Failure-driven discovery; merged libraries beat single runs
Feb 2026 SkillsBench 27 87 tasks, 18 model-harness configs 33.9% → 50.5% with curated skills (+16.6 pp); focused skills beat large bundles Human-curated skills help on average

Three limits apply to nearly every row. Most gains are measured in distribution: skills are learned and tested on the same benchmark family, and only EvoSkill's BrowseComp transfer, SkillOpt's cross-benchmark check, and Code2Skill's breadth provide out-of-distribution evidence. Few papers report a matched-budget baseline that gives the no-skill agent the extra tokens and calls the skill loop consumed (see 14, Measuring self-improvement), and ContinualSkillBench suggests that baseline would absorb much of the gain. Most are single-group preprints; the positive results have not been independently replicated at the scale the negative results in the next section have.

Section 06 / 09

#Where it breaks at library scale

The skill loop works. The failures appear when the library gets large, shared, and old, and they are failures of activation, acceptance, and maintenance, not of skill writing.

Relevance is the hazard

The 307-failure study is a forensic account of how obedient models misuse good advice 1. Excessive Verification and Heavy Implementation Pipeline are not bugs in the skills; they are mismatches between a skill's generality and a task's needs. A skill written to be safe across many tasks includes checks and scaffolding that most individual tasks do not need, and a model trained to follow instructions treats "consider validating X" as "validate X". Because 182 of the 307 failures were tasks that passed at a higher cost, a with/without eval that scores only pass rate would record them as neutral or positive.

The first fixes put a gate between retrieval and execution. RADEG (August 2026) learns a cheap surrogate that predicts whether running a retrieved skill bundle is worth it, trained on matched rollouts in which one skill of the bundle is deleted, added, or replaced so that the effect of composition on verifier reward is isolated; on 288 held-out rollouts it beat relevance-based and random gating across execution budgets 2. SkillApt (September 2026) makes the same move per skill, deciding LOAD or ABSTAIN from matched with/without runs in similar historical states. On its SRA-Bench evaluation it matched BM25 top-1 accuracy (0.838 against 0.838) while activating skills 31.5% of the time instead of 100% and cutting mean token use by 74.3% 3. Its authors also found that both skill utility and how learnable the activation boundary is vary across base models 3. How a skill is invoked matters too: a September 2026 study found that running skill packages as subagents in fresh context windows beat loading them into the main context when the packages expose clear input–output contracts, at the cost of coordination tokens 28.

Fig 08.1 · Relevance and activation, gated apartSkillApt · Wix Helpmate · 2026
RELEVANCE Is this skill on topic? ACTIVATION Should it run now? RESULT SKILLAPT SRA-Bench BM25 top-1 retrieval LOAD or ABSTAIN from with/without runs Accuracy 0.838, unchanged Activation 100% → 31.5% Mean tokens −74.3% WIX HELPMATE 756.6K messages Semantic matcher Deterministic gate hard-stop conditions 59.4% of matched pairs removed Skill-description context −90.5% REPLAY · EVERY SKILL EXPOSED 78 of 1,000 conversations picked a skill production would block
Retrieval finds skills; a gate decides. SkillApt decides LOAD or ABSTAIN for each retrieved skill from matched with/without runs in similar historical states, and on SRA-Bench matched BM25 top-1 accuracy while activating skills 31.5% of the time 3. Wix's production pipeline puts a deterministic hard-stop gate after semantic matching; when every skill was exposed in a replay, the model picked a blocked skill in 78 of 1,000 conversations 35. The two lanes come from different systems and are not one experiment.

Effects cancel in aggregate

If relevant skills sometimes hurt, the natural remedy is to measure each skill and prune the harmful ones. ASSAY (June 2026) shows why an aggregate measurement cannot do that 29. Using randomized masking (running tasks with random subsets of the library switched off and estimating each skill's causal contribution), it found "pervasive causal heterogeneity": individual skills "routinely help on some task types while hurting on others, yet their opposing effects cancel in aggregate, making them invisible to global curation methods" 29. A skill with zero net effect may be very good for one task type and quite bad for another; global pruning keeps it or drops it and loses either way. ASSAY's per-task suppression beat prior curation methods across seven base models from four providers on AppWorld and tau-bench 29, but randomized masking needs many runs per skill per task type, and an ecosystem registry holds tens of thousands of skills. SkillApt and RADEG are cheaper relatives of the same idea: they estimate a skill's effect in the current state instead of on average.

The acceptor commits noise

Every skill system above keeps an edit because some score went up. PACE (June 2026) formalized why that decision is unreliable when repeated 30. Applying "keep if the score improved" hundreds of times against a noisy development set is adaptive multiple testing with no correction. On Qwen2.5 agents from 0.5B to 3B parameters, greedy acceptance committed 30–42% false edits and 10–33% harmful ones when real gains were available; when none were, 72–100% of commits were false (13–21 spurious self-modifications per run), and the most fragile agent lost 4.9 points 30. PACE replaces the greedy rule with an e-process, a sequential test that accumulates paired evidence for candidate over incumbent on identical instances and commits only past a threshold; false commits fell to about zero at roughly 18% lower evaluation cost 30. Report 06 (Prompt and program optimization) owns PACE in depth.

PACE's experiments are on prompt edits in small models, not skill libraries, so the specific rates do not carry over. The structure does. SkillOpt, EvoSkill, SkillRevise, and the September optimizers all use some form of keep-if-better, and COBRA-Skills' bandit makes evaluation cheaper without changing the commit rule 16. A skill commit is a hypothesis test, and most skill loops run it without error control.

Skills go stale in silence

The last failure is time. Repo2Skill-Evo (August 2026) cast each software release as a skill-maintenance task: given skills written for version 1 and the patch to version 2, update what is obsolete and keep what still holds 25. Across 57 real repositories and 105 release transitions, every transition invalidated part of the version-1 skill set, and six frontier agents reached only 29.9%–69.7% avg@3 macro F1 at the repair 25. The errors split two ways: leaving stale content untouched, or editing too broadly and deleting guidance that was still correct.

A stale skill does not throw an error. It gives confident, specific, outdated instructions, and because it is relevant, the agent follows them. The 307-failure finding and the staleness finding compound: the skills most likely to be loaded are the ones most likely to mislead when the code under them changes. In practice, humans still do the maintenance. A September 2026 longitudinal study mined five public AI-skill repositories (873 commits, 143 skill files, 254 substantive post-creation edits from October 2025 to June 2026) and found that every substantive edit was authored or merged through a named human account, while 62% carried an AI co-author trailer 31. Most reuse gets no maintenance at all: a July 2026 study of 18,463 registry skills and 23,199 skills across 5,876 GitHub repositories reports that 53% of reused skills were never modified after adoption 32. Public skill maintenance is a human-governed, AI-assisted loop, and much of it never runs.

Section 07 / 09

#State of play, 2026

The author's June post described a field moving from authoring to operations 10. Five things changed between July and September 2026.

Activation split from retrieval. The 307-failure taxonomy (August) located harm in relevant skills 1, and RADEG (August) and SkillApt (September) turned "should this skill run now?" into a learned decision with its own evidence 2,3. A production system reached the same design independently (see 16, What ships).

Recursion reached the improvement procedure, and the proposer got cheaper. MetaSkill-Evolve (July) is the first skill system whose improver evolves itself 23; COBRA-Skills and GraphSkillEvo (both September) beat SkillOpt, and SkillAdam (September) stabilizes the edit loop 16,17,18. None of them pairs its proposer with a statistically controlled acceptor.

The evidence that explicit skills are the active ingredient weakened. ContinualSkillBench (August) found in-context learning comparable to explicit skill maintenance on average 19, and Code2Skill (September) showed that a million skills mined from code, with no agent experience, beat trajectory-derived banks 11. What an agent learns by doing is not obviously better than what it can read.

The line between skills and weights blurred from both sides. SKT (August) synthesized 27,164 verified trajectories from 2,000 public skills and fine-tuned models on them, which consistently improved skill use across benchmarks and harnesses 33. A skill-internalization subfield now trains skills into weights and removes them from context; report 12 (Consolidation and co-evolution) covers it. Prime Intellect's Prime Agent (August) exposes prompts, skills, memory, and sub-agents to the agent through one create/read/update/delete interface driven by its own trajectories 34, so the boundary between "skill", "tool", and "harness" is thinning in practice (see 09, Harness self-improvement).

Security research caught up with self-authoring. Earlier skill-security work studied malicious skills that attackers submit. August and September 2026 papers showed attacks that ride the self-improvement loop itself, covered in the next section.

Section 08 / 09

#What ships

The shipped version of the self-extending agent is open source and largely human-reviewed; report 16 (What ships) covers the products in depth.

The clearest production evidence on skill selection comes from customer care. In August 2026, Wix described the selection pipeline for Helpmate, its customer-care assistant: a semantic matcher, then a deterministic gate that removes skills whose hard-stop conditions already hold for the account, then the LLM's choice 35. Across 756.6K user messages, the gate removed 59.4% of skill–message pairs that survived semantic matching, and together the two stages cut skill-description context by 90.5% against exposing all ten skills. In a replay of 1,000 risk-enriched conversations with every skill exposed, the model picked a skill production would have blocked in 78 of them (7.8%) 35. A relevant but non-executable skill steers the model when it is visible, which is the 307-failure mechanism measured in production.

OpenClaw (390,396 GitHub stars on September 24, 2026) treats skills as packages: a CLI to install, search, update, and publish, per-agent allowlists, and a public registry, ClawHub, which lists about 67,300 skills 36,37. A September 2026 study of three ClawHub snapshots describes a boom that has crested: the observable stock nearly doubled in 91 days in the first half of 2026, and monthly listing creation was falling from its spring peak by the end of the window 24. Attention is concentrated (the top 10% of skills received 46.93% of downloads) and scrutiny is thin: 77.86% of skills have zero stars and zero comments, while 85.06% of readable skills carry evidence of privileged actions 24. Nous Research's Hermes Agent (248,668 stars, MIT license) builds the self-improvement loop into the product: its README describes "a built-in learning loop" that "creates skills from experience, improves them during use, nudges itself to persist knowledge, searches its own past conversations, and builds a deepening model of who you are across sessions" 38. The README describes what the loop writes; it does not describe a test that gates what gets kept.

The same shape that made skills portable made them a supply chain. Between February and May 2026, the ClawHavoc campaign pushed malicious skills onto ClawHub; a Cloud Security Alliance research note on June 24, 2026 put the confirmed count at 1,184 39. Code scanners are not blind to skill malware in general (VirusTotal detected 72.8% of confirmed malicious skills that bundled code 40), but the 0% detection figure that circulated applies to payload-less Semantic Compliance Hijacking, where the attack lives entirely in natural-language instructions 41. Scanners also disagree with each other: on the 61,990 ClawHub skills that three scanners all cover, they disagreed on 23,702, and after human adjudication their weighted sensitivity ranged from 21.67% to 61.06% 24. In September 2026, SkillShift showed a subtler class: skills that keep their declared task and valid outputs while covertly steering decisions, reaching attacker-favored selection rates of 81.33% in agentic commerce and 63.33% in software-dependency choices with utility fully preserved, undetected by the scanners evaluated 42. Report 15 (Failure modes and safety) covers the incidents and defenses.

The attack that matters most for self-improvement is self-poisoning. EVOMAL (August 2026) plants malicious skills in a shared library without ever invoking them; a self-evolving coding agent that writes new tools by imitating retrieved skills copies the payload into its own new skill, stores it, and runs it 26. Across six models on 153 SWE-bench Verified tasks, 20.3%–41.8% of tasks added a newly authored malicious skill, and poisoned libraries ended up holding 4.9–9.0 times as many malicious skills as were planted. After the planted skills were removed, Qwen3 still showed a 68% self-poisoning rate at round five because its own copies remained 26. Those copies evade defenses that check attacker-submitted names, code, and signatures, because the agent wrote them. A counter-prompt that discourages template-style copying cut the rate to at most 6.7% without significant loss in task completion 26. Package provenance and signed catalogs secure installation; a library the agent writes for itself needs its own provenance, down to which retrieved skill each new one imitated.

Section 09 / 09

#Open problems

Staleness detection without release signals. Repo2Skill-Evo gives agents the patch and asks them to repair the skills; even then the best agent reaches 69.7% F1 25. In deployment, most drift arrives without a clean diff: an API changes behavior, a website redesigns, a dependency deprecates a flag. Code2Skill-scale banks pin a million records to moving code 11. Detecting that a skill has gone stale, before it misleads a run, is unsolved, and the harmful-skills result means the cost of missing it concentrates in the skills used most.

Per-task causal attribution at scale. ASSAY shows that the right unit of evaluation is skill × task type, not skill 29, and SkillApt shows the effect can be estimated per state from matched runs 3. Both need with/without evidence that is expensive to collect for a registry of tens of thousands of skills. A cheap estimator, learned from routine logs, would let per-state activation move from benchmarks into registries. Without one, every curation decision averages away the signal that matters.

Convergence and gating of meta-improvement. MetaSkill-Evolve shows that evolving the improvement procedure helps, by 23.54 points on one benchmark and 1.92 on another 23. Nobody knows whether the two-timescale process converges, oscillates, or overfits its procedure to the benchmark that drives it, or how a meta-level edit should be gated when its effect shows up only after many task-level cycles. PACE-style sequential testing for task-level commits exists; the equivalent for procedure-level commits does not.

Provenance for self-authored skills. EVOMAL's worm survives because nothing records that a new skill was modeled on a retrieved one 26. Lineage tracking for agent-written skills, and acceptance tests that check behavior rather than signatures, are missing from every system covered here.

Three conclusions survive the caveats. Skills are the most productive non-parametric self-improvement mechanism in 2026: they compound within a domain, transfer across models, and ship in every major coding agent. The proposer side is no longer the bottleneck: many methods write helpful skills, from traces or straight from code, and the September optimizers made it cheaper still. The binding constraints are activation, acceptance, and maintenance: skills that are relevant but over-applied, effects that vary by task and cancel in aggregate, libraries that rot at every release, and a library the agent writes for itself that an attacker can also write to. The agent that extends itself well will be the one that knows which of its own skills to leave unloaded.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]Dong et al., “Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents,” arXiv:2608.11888, August 2026. https://arxiv.org/abs/2608.11888
  2. [2]He, Wen, Gu, Li et al., “From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents” (RADEG), arXiv:2608.09168, August 10, 2026. https://arxiv.org/abs/2608.09168
  3. [3]Shuang Guo, “SkillApt: Learning When to Activate Agent Skills from Counterfactual Evidence,” arXiv:2609.26863, September 22, 2026. https://arxiv.org/abs/2609.26863
  4. [4]Wang et al., “Voyager: An Open-Ended Embodied Agent with Large Language Models,” NeurIPS 2023 (arXiv:2305.16291, May 2023). https://arxiv.org/abs/2305.16291
  5. [5]Cai et al., “Large Language Models as Tool Makers,” ICLR 2024 (arXiv:2305.17126, May 2023). https://arxiv.org/abs/2305.17126
  6. [6]“SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills,” arXiv:2504.07079, April 2025. https://arxiv.org/abs/2504.07079
  7. [7]Qin et al., “Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution,” arXiv:2505.20286, May 2025. https://arxiv.org/abs/2505.20286
  8. [8]Anthropic, “Equipping agents for the real world with Agent Skills,” Anthropic Engineering, October 16, 2025. https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
  9. [9]Agent Skills open specification, agentskills.io, published as an open standard December 2025. https://agentskills.io
  10. [10]Soham Shah, “Skill Libraries Need CI, Not More Prompts,” sohams.world, June 9, 2026 (sources therein for Codex skills docs and the GitHub `gh skill` changelog of April 16, 2026). https://sohams.world/posts/skill-evolution-landscape/
  11. [11]Tong, Wang, Wang, Li et al., “Grounded Skill Synthesis from Code at Scale for Agentic Intelligence” (Code2Skill), arXiv:2609.05571, September 4, 2026. https://arxiv.org/abs/2609.05571
  12. [12]Ni et al., “Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills,” arXiv:2603.25158, March 2026 (v5 June 2026). https://arxiv.org/abs/2603.25158
  13. [13]Alzubi et al., “EvoSkill: Automated Skill Discovery for Multi-Agent Systems,” arXiv:2603.02766, March 2026. https://arxiv.org/abs/2603.02766
  14. [14]Liu et al., “SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision,” arXiv:2606.01139, May 2026 (v4 September 2026). https://arxiv.org/abs/2606.01139
  15. [15]Yang et al. (Microsoft), “SkillOpt: Executive Strategy for Self-Evolving Agent Skills,” arXiv:2605.23904, May 2026. https://arxiv.org/abs/2605.23904
  16. [16]Lu, Wang, Li, Mao et al., “COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization,” arXiv:2609.11682, September 10, 2026. https://arxiv.org/abs/2609.11682
  17. [17]Sun, Zheng, Wang, Lu, “GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills,” arXiv:2609.21749, September 18, 2026. https://arxiv.org/abs/2609.21749
  18. [18]Li, Fan, Liu, Zhang et al., “SkillAdam: Stable and Efficient Skill Evolution for Agents,” arXiv:2609.08944, September 8, 2026. https://arxiv.org/abs/2609.08944
  19. [19]Guan, Wang, Yang, Cao et al., “ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?,” arXiv:2608.03874, August 4, 2026. https://arxiv.org/abs/2608.03874
  20. [20]Zhong et al., “SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation,” COLM 2026 (arXiv:2604.20087, April 2026). https://arxiv.org/abs/2604.20087
  21. [21]“From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills,” arXiv:2605.23899, May 2026. https://arxiv.org/abs/2605.23899
  22. [22]Zhang et al., “CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification,” COLM 2026 (arXiv:2604.01687, April 2026). https://arxiv.org/abs/2604.01687
  23. [23]Wang et al., “MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution,” arXiv:2607.05297, July 2026. https://arxiv.org/abs/2607.05297
  24. [24]Yunpeng Xiong, Ting Zhang, “After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem,” arXiv:2609.17274, September 15, 2026. https://arxiv.org/abs/2609.17274
  25. [25]Zhang et al., “Repo2Skill-Evo: Repository Skills Go Stale in Silence,” arXiv:2608.21964, August 2026. https://arxiv.org/abs/2608.21964
  26. [26]Wu, Shi, Li, Zhao et al., “EVOMAL: Self-Poisoning in Self-Evolving Coding Agents,” arXiv:2608.25776, August 26, 2026. https://arxiv.org/abs/2608.25776
  27. [27]Li et al., “SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks,” arXiv:2602.12670, February 2026 (v4 June 2026). https://arxiv.org/abs/2602.12670
  28. [28]Piriyakulkij, Lawrence, Curth, Karmalkar et al., “Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks,” arXiv:2609.09233, September 7, 2026. https://arxiv.org/abs/2609.09233
  29. [29]“Not All Skills Help: Measuring and Repairing Agent Knowledge” (ASSAY), arXiv:2606.15390, June 2026. https://arxiv.org/abs/2606.15390
  30. [30]Zayx Shawn, “PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents,” arXiv:2606.08106, June 2026. https://arxiv.org/abs/2606.08106
  31. [31]Chen Shen, Estevam Hruschka, “Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance,” arXiv:2609.05677, September 4, 2026. https://arxiv.org/abs/2609.05677
  32. [32]“From Registry to Repository: How AI Agent Skills Are Written, Adapted, and Maintained,” arXiv:2607.00911, July 2026. https://arxiv.org/abs/2607.00911
  33. [33]“SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation,” arXiv:2608.02287, August 2026. https://arxiv.org/abs/2608.02287
  34. [34]Prime Intellect, “Prime Agent” (Continual Harness), arXiv:2608.23552, August 2026; blog https://www.primeintellect.ai/blog/prime-agent. https://arxiv.org/abs/2608.23552
  35. [35]Ashkenazi, Kloz, Ulianchenko (Wix), “Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale,” arXiv:2608.01050, August 2, 2026. https://arxiv.org/abs/2608.01050
  36. [36]OpenClaw, GitHub repository (star count as of September 24, 2026) and skills documentation. https://github.com/openclaw/openclaw ; https://docs.openclaw.ai/tools/skills
  37. [37]ClawHub skill registry, accessed September 24, 2026. https://hub.openclaw.ai/skills
  38. [38]Nous Research, Hermes Agent, GitHub repository and README (star count as of September 24, 2026). https://github.com/NousResearch/hermes-agent
  39. [39]Cloud Security Alliance, “AI skill supply-chain attacks” research note, June 24, 2026. https://labs.cloudsecurityalliance.org/wp-content/uploads/2026/06/CSA_research_note_ai-skill-supply-chain-attacks_20260624-csa-styled.pdf
  40. [40]arXiv:2606.01494 (malicious agent-skill detection study; VirusTotal detection of code-bundled malicious skills), June 2026. https://arxiv.org/abs/2606.01494
  41. [41]arXiv:2605.11418 (Semantic Compliance Hijacking: payload-less skill attacks), May 2026. https://arxiv.org/abs/2605.11418
  42. [42]Li, Chen, Zhou, Pu et al., “A Finger on the Scale: Covert Policy Steering through Agentic Skills” (SkillShift), arXiv:2609.02564, September 2, 2026. https://arxiv.org/abs/2609.02564