#When the right skill makes the agent worse
In August 2026, Dong et al. ran agents on SkillsBench and SWE-Skills-Bench twice per task: once with a loaded skill, once without it or with a semantically matched reference skill. Wherever the skill-guided run did worse, they traced the regression to the skill. They found 307 skill-induced failures: 125 functional failures, where the task broke, and 182 efficiency regressions, where the task succeeded at a higher cost 1. Irrelevant skills were not the main culprit. Seemingly relevant skills caused most of the functional failures.
The efficiency taxonomy shows how. The largest category, Excessive Verification, accounts for 67 cases; the next, Heavy Implementation Pipeline, accounts for 30 1. In both, the skill was on topic and its advice was reasonable in general. The agent read a validation checklist as a required procedure and ran every check, or read a construction recipe (subprocess workflows, multi-stage conversion, runtime simulation) as the only acceptable path and built all of it for a task that needed a fraction. In the authors' words, skills "often turn validation checklists and construction recipes into mandatory work" 1. An irrelevant skill gets ignored. A relevant skill gets obeyed, including the parts that do not apply.
That result inverts the obvious failure model for skill libraries. The obvious worry is retrieval: load the wrong skill and the agent goes astray, so improve retrieval. The 307-failure study says a large share of the damage comes from loading the right skill and following it too literally. Better retrieval makes that worse, because it loads more relevant skills more often. Within six weeks, two papers answered with the same design move: treat "is this skill relevant?" and "should this skill run now?" as separate decisions, and learn the second one from matched with/without evidence 2,3.
Library-scale failures of this kind are a 2026 phenomenon. Before that, the skills literature was about getting the loop to work at all: can an agent write reusable procedures, verify them, and get better by reusing them? It can, and the way that loop was built explains why its failures showed up only once libraries got large, shared, and long-lived.
#How the self-improvement loop works
Every self-extending agent runs the same loop: gather evidence, propose a new or revised skill, test it, commit or reject, and retrieve it later. The systems differ in what they mine, how they propose, and what decides.
Mining code before experience: Code2Skill
The newest extraction work skips agent experience altogether. In September 2026, Code2Skill turned code units from 19,769 popular, actively maintained GitHub repositories into skill records (atomic operations, composite workflows, recurring patterns), verified each one by reconstructing it without seeing the source body and then comparing against the source, and kept 1,006,822 accepted records with provenance metadata 11. Across 72 protocol-matched evaluations covering nine model settings and eight benchmarks, retrieving from that bank improved models by 11.7% on average over matched baselines and beat them in 57 cases; under a shared interface it also beat trajectory-derived skill banks on all seven shared benchmarks 11. Repository-derived skills solve the cold start that trace mining cannot, because they exist before the agent has done anything. They also pin a million records to code that will change, which makes staleness (see Where it breaks) more pressing.
Consolidating many traces at once: Trace2Skill
Trace mining still carries the strongest transfer evidence. Early extraction was sequential: finish an episode, write or edit a skill, move on, which overfits to whichever episode came last. Trace2Skill (March 2026, v5 June 2026) consolidates many trajectories in parallel instead, inducing recurring patterns across a broad set of traces, resolving conflicts between them, and writing a unified skill directory; it outperforms both sequential skill editing and ReasoningBank-style retrieval of raw experience 12.
Its most important result concerns transfer. Skills that Qwen3.5-35B evolved on its own trajectories improved a Qwen3.5-122B agent by up to 57.65 absolute percentage points on WikiTableQuestions, with no weight update 12. A smaller model's procedural lessons, written as files, made a larger model much better at the same domain. The comparison is within a model family, "up to" marks a best case, and the paper is labeled work in progress; within those limits, it is the strongest evidence in the series that a lesson written into text survives a change of model. That is the portability property the series keeps returning to (see 12, Consolidation and co-evolution): code and structured procedure transfer across models where tuned prose often does not.
Mining failures: EvoSkill and SkillRevise
Successful traces show one path that worked. Failed traces show where the agent's default behavior breaks, which is what a skill should fix. EvoSkill (Sentient AI and Virginia Tech, March 2026) analyzes execution failures, proposes new skills or edits, and keeps only those on a Pareto frontier of held-out validation performance, with the model frozen throughout 13. On OfficeQA, exact match rose from 60.6% to 67.9%; on SealQA, from 26.6% to 38.7%; and skills learned on SealQA transferred zero-shot to BrowseComp for a 5.3% gain 13. Merging libraries from independent runs beat any single run (67.9% against 64.5%), and gains flattened as the training fraction grew 13. SkillRevise (HKUST, May 2026, v4 September 2026) handles the cold start with one imperfect skill: it diagnoses defects from execution evidence, retrieves general repair principles, applies edits anchored to the observed failure, and keeps the best re-executed candidate; on SkillsBench, base-agent success rose from 36.05% to 61.63% 14.
The optimizer race: SkillOpt and its September challengers
SkillOpt (Microsoft, May 2026) treats skill text as trainable external state 15. It proposes bounded edits, keeps a history of rejected edits so it does not cycle back to them, and commits an edit only when held-out validation improves. Across 6 benchmarks, 7 target models, and 3 harnesses (direct chat, Codex, Claude Code), it was best or tied in all 52 cells; with GPT-5.5 it added 23.5 points in direct chat, 24.8 inside Codex, and 19.1 inside Claude Code 15.
SkillOpt became the baseline to beat within four months. In September 2026, COBRA-Skills cast skill optimization as budgeted sequential search, using a contextual bandit to decide which candidates deserve an expensive evaluation; it reported the strongest average across six benchmarks and three target models at 55–58% lower optimization cost than SkillOpt, using 50 optimization examples per benchmark 16. GraphSkillEvo represented a skill as a graph of steps and transitions and evolved a population with mutation and crossover, beating SkillOpt by 4.01% average accuracy on GPT-5.4-nano and 1.76% on GPT-5.4 across five benchmarks 17. SkillAdam added an optimization memory and a volatility-driven edit budget, functional analogues of Adam's two moments, so that iteration-local feedback stops overwriting earlier corrections 18. The proposer is getting cheaper, more structured, and more stable. None of the three changes the acceptance rule: an edit survives because a validation score went up.
Who judges the skill
A skill loop is bounded by its evaluator, and the August 2026 evidence questions whether explicit skills are doing the work at all. ContinualSkillBench built five domains of 100 interconnected subtasks each, ordered by difficulty with chances for cross-task reuse 19. Sequential execution generally improved performance, but plain in-context learning performed comparably to explicit skill maintenance on average, "suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone"; explicit skills helped selectively on tasks needing reusable procedures or precise outputs, and less capable models accumulated larger, more fragmented collections of task-specific skills 19. The result echoes CL-Bench's finding for memory systems (see 07, Experiential memory, and 14, Measuring self-improvement).
Earlier 2026 benchmarks point the same way. SkillLearnBench (CMU and Amazon AGI; COLM 2026) found that every continual skill-learning method beat the no-skill baseline, but none led consistently, stronger backbones did not reliably produce better skills, and self-feedback alone induced recursive drift, while repeated external feedback produced real improvement 20. A May 2026 lifecycle study found that a model can be a strong skill extractor and a weak skill consumer, and that model-generated skills show non-trivial negative transfer even when they help on average 21. Where no ground-truth tests exist, CoEvoSkills (COLM 2026) evolves a Surrogate Verifier alongside the Skill Generator without showing it test content, and beat five baselines on SkillsBench for both Claude Code and Codex 22. The ordering from the series argument holds: external checks beat co-evolved judges, which beat self-judgment.
#Improving the improver
Every system so far has a fixed improvement procedure written by its authors: how to analyze a failure, which traces to retrieve, how much budget to spend, how to phrase an edit. MetaSkill-Evolve (July 2026) makes that procedure a skill too 23. Its improvement pipeline has five roles (Analyzer, Retriever, Allocator, Proposer, Evolver), and a meta-skill encodes how each should operate. Task skills evolve on a fast loop from execution traces. The meta-skill evolves on a slower loop, "under the same pipeline applied to itself, with no additional model or objective" 23. All five roles run on the same frozen backbone.
Over the raw backbone, MetaSkill-Evolve added 23.54 points on OfficeQA, 16.09 on SealQA, and 1.92 on ALFWorld, and it beat no-skill, static-skill, and single-level evolution baselines on all three 23. Beating single-level evolution is the result that matters: it says evolving the procedure adds something over evolving the skills with a fixed procedure.
The recursion is real and bounded. The meta-skill improves the improvement procedure, and it is improved by that same procedure; there is no third level, and nothing in the paper shows the two-level process converging or continuing to improve beyond the reported runs. The spread across benchmarks (23.54 points on OfficeQA, 1.92 on ALFWorld) suggests the meta-level helps most where the base skill loop has the most room. Two timescales also recur across the series: explore fast in context, consolidate slowly. Here the slow store is still text, a meta-skill rather than weights, so the whole recursive stack stays inspectable and can be rolled back. Report 13 (Recursive self-improvement) covers the definitional debate about whether this counts as recursive self-improvement.
#What the evidence shows
Across the 2026 literature, gains are large when skills target a specific domain and are checked by execution, and smaller and more variable when the skill set is broad or the check is weak.
| Date | System | Setting | Baseline → result | What it demonstrates |
|---|---|---|---|---|
| Sep 2026 | SkillApt 3 | SRA-Bench | Same accuracy as BM25 top-1 (0.838 vs 0.838); skill activation 100% → 31.5%; mean tokens −74.3% | Deciding whether to load a retrieved skill is a separate, learnable decision |
| Sep 2026 | Code2Skill 11 | 72 evaluations, 9 model settings, 8 benchmarks | +11.7% average over matched baselines; better in 57/72; beats trajectory-derived banks on 7/7 | Skills mined from 19,769 repos work before any agent experience |
| Sep 2026 | COBRA-Skills 16 | 6 benchmarks, 3 target models | Strongest average among compared methods at 55–58% lower optimization cost than SkillOpt | The proposer side is getting cheaper |
| Sep 2026 | OpenClaw registry study 24 | 61,990 skills covered by three scanners | Scanners disagree on 23,702; weighted sensitivity 21.67%–61.06% | Registry security cannot rest on one scanner |
| Aug 2026 | Repo2Skill-Evo 25 | 57 repos, 105 release transitions | Every transition invalidates part of the skills; frontier agents 29.9%–69.7% macro F1 at maintenance | Skills decay with every release |
| Aug 2026 | EVOMAL 26 | 6 models, 153 SWE-bench Verified tasks | 20.3%–41.8% of tasks add a new malicious skill; poisoned libraries hold 4.9–9.0× the planted count | Self-authoring turns one planted skill into a worm |
| Aug 2026 | Harmful-skills study 1 | SkillsBench, SWE-Skills-Bench | 307 skill-induced failures (125 functional, 182 efficiency) | Relevant skills cause most failures |
| Aug 2026 | ContinualSkillBench 19 | 5 domains × 100 subtasks | In-context learning comparable to explicit skill maintenance on average | Much apparent skill learning is context adaptation |
| Jul 2026 | MetaSkill-Evolve 23 | OfficeQA / SealQA / ALFWorld | +23.54 / +16.09 / +1.92 pp over raw backbone | Evolving the improvement procedure adds over single-level evolution |
| May 2026 | SkillOpt 15 | 52 model/benchmark/harness cells | Best or tied in 52/52; GPT-5.5 +23.5 (chat), +24.8 (Codex), +19.1 (Claude Code) | Gated edit loop beats other optimizers and human skills |
| May 2026 | SkillRevise 14 | SkillsBench | 36.05% → 61.63% base-agent success | Execution-anchored repair from a cold start |
| Mar 2026 | Trace2Skill 12 | WikiTableQuestions | Qwen3.5-35B-evolved skills lift Qwen3.5-122B by up to 57.65 pp | Skills transfer across model scale without weight updates |
| Mar 2026 | EvoSkill 13 | OfficeQA; SealQA → BrowseComp | 60.6% → 67.9%; 26.6% → 38.7%; +5.3% zero-shot on BrowseComp | Failure-driven discovery; merged libraries beat single runs |
| Feb 2026 | SkillsBench 27 | 87 tasks, 18 model-harness configs | 33.9% → 50.5% with curated skills (+16.6 pp); focused skills beat large bundles | Human-curated skills help on average |
Three limits apply to nearly every row. Most gains are measured in distribution: skills are learned and tested on the same benchmark family, and only EvoSkill's BrowseComp transfer, SkillOpt's cross-benchmark check, and Code2Skill's breadth provide out-of-distribution evidence. Few papers report a matched-budget baseline that gives the no-skill agent the extra tokens and calls the skill loop consumed (see 14, Measuring self-improvement), and ContinualSkillBench suggests that baseline would absorb much of the gain. Most are single-group preprints; the positive results have not been independently replicated at the scale the negative results in the next section have.
#Where it breaks at library scale
The skill loop works. The failures appear when the library gets large, shared, and old, and they are failures of activation, acceptance, and maintenance, not of skill writing.
Relevance is the hazard
The 307-failure study is a forensic account of how obedient models misuse good advice 1. Excessive Verification and Heavy Implementation Pipeline are not bugs in the skills; they are mismatches between a skill's generality and a task's needs. A skill written to be safe across many tasks includes checks and scaffolding that most individual tasks do not need, and a model trained to follow instructions treats "consider validating X" as "validate X". Because 182 of the 307 failures were tasks that passed at a higher cost, a with/without eval that scores only pass rate would record them as neutral or positive.
The first fixes put a gate between retrieval and execution. RADEG (August 2026) learns a cheap surrogate that predicts whether running a retrieved skill bundle is worth it, trained on matched rollouts in which one skill of the bundle is deleted, added, or replaced so that the effect of composition on verifier reward is isolated; on 288 held-out rollouts it beat relevance-based and random gating across execution budgets 2. SkillApt (September 2026) makes the same move per skill, deciding LOAD or ABSTAIN from matched with/without runs in similar historical states. On its SRA-Bench evaluation it matched BM25 top-1 accuracy (0.838 against 0.838) while activating skills 31.5% of the time instead of 100% and cutting mean token use by 74.3% 3. Its authors also found that both skill utility and how learnable the activation boundary is vary across base models 3. How a skill is invoked matters too: a September 2026 study found that running skill packages as subagents in fresh context windows beat loading them into the main context when the packages expose clear input–output contracts, at the cost of coordination tokens 28.
Effects cancel in aggregate
If relevant skills sometimes hurt, the natural remedy is to measure each skill and prune the harmful ones. ASSAY (June 2026) shows why an aggregate measurement cannot do that 29. Using randomized masking (running tasks with random subsets of the library switched off and estimating each skill's causal contribution), it found "pervasive causal heterogeneity": individual skills "routinely help on some task types while hurting on others, yet their opposing effects cancel in aggregate, making them invisible to global curation methods" 29. A skill with zero net effect may be very good for one task type and quite bad for another; global pruning keeps it or drops it and loses either way. ASSAY's per-task suppression beat prior curation methods across seven base models from four providers on AppWorld and tau-bench 29, but randomized masking needs many runs per skill per task type, and an ecosystem registry holds tens of thousands of skills. SkillApt and RADEG are cheaper relatives of the same idea: they estimate a skill's effect in the current state instead of on average.
The acceptor commits noise
Every skill system above keeps an edit because some score went up. PACE (June 2026) formalized why that decision is unreliable when repeated 30. Applying "keep if the score improved" hundreds of times against a noisy development set is adaptive multiple testing with no correction. On Qwen2.5 agents from 0.5B to 3B parameters, greedy acceptance committed 30–42% false edits and 10–33% harmful ones when real gains were available; when none were, 72–100% of commits were false (13–21 spurious self-modifications per run), and the most fragile agent lost 4.9 points 30. PACE replaces the greedy rule with an e-process, a sequential test that accumulates paired evidence for candidate over incumbent on identical instances and commits only past a threshold; false commits fell to about zero at roughly 18% lower evaluation cost 30. Report 06 (Prompt and program optimization) owns PACE in depth.
PACE's experiments are on prompt edits in small models, not skill libraries, so the specific rates do not carry over. The structure does. SkillOpt, EvoSkill, SkillRevise, and the September optimizers all use some form of keep-if-better, and COBRA-Skills' bandit makes evaluation cheaper without changing the commit rule 16. A skill commit is a hypothesis test, and most skill loops run it without error control.
Skills go stale in silence
The last failure is time. Repo2Skill-Evo (August 2026) cast each software release as a skill-maintenance task: given skills written for version 1 and the patch to version 2, update what is obsolete and keep what still holds 25. Across 57 real repositories and 105 release transitions, every transition invalidated part of the version-1 skill set, and six frontier agents reached only 29.9%–69.7% avg@3 macro F1 at the repair 25. The errors split two ways: leaving stale content untouched, or editing too broadly and deleting guidance that was still correct.
A stale skill does not throw an error. It gives confident, specific, outdated instructions, and because it is relevant, the agent follows them. The 307-failure finding and the staleness finding compound: the skills most likely to be loaded are the ones most likely to mislead when the code under them changes. In practice, humans still do the maintenance. A September 2026 longitudinal study mined five public AI-skill repositories (873 commits, 143 skill files, 254 substantive post-creation edits from October 2025 to June 2026) and found that every substantive edit was authored or merged through a named human account, while 62% carried an AI co-author trailer 31. Most reuse gets no maintenance at all: a July 2026 study of 18,463 registry skills and 23,199 skills across 5,876 GitHub repositories reports that 53% of reused skills were never modified after adoption 32. Public skill maintenance is a human-governed, AI-assisted loop, and much of it never runs.
#State of play, 2026
The author's June post described a field moving from authoring to operations 10. Five things changed between July and September 2026.
Activation split from retrieval. The 307-failure taxonomy (August) located harm in relevant skills 1, and RADEG (August) and SkillApt (September) turned "should this skill run now?" into a learned decision with its own evidence 2,3. A production system reached the same design independently (see 16, What ships).
Recursion reached the improvement procedure, and the proposer got cheaper. MetaSkill-Evolve (July) is the first skill system whose improver evolves itself 23; COBRA-Skills and GraphSkillEvo (both September) beat SkillOpt, and SkillAdam (September) stabilizes the edit loop 16,17,18. None of them pairs its proposer with a statistically controlled acceptor.
The evidence that explicit skills are the active ingredient weakened. ContinualSkillBench (August) found in-context learning comparable to explicit skill maintenance on average 19, and Code2Skill (September) showed that a million skills mined from code, with no agent experience, beat trajectory-derived banks 11. What an agent learns by doing is not obviously better than what it can read.
The line between skills and weights blurred from both sides. SKT (August) synthesized 27,164 verified trajectories from 2,000 public skills and fine-tuned models on them, which consistently improved skill use across benchmarks and harnesses 33. A skill-internalization subfield now trains skills into weights and removes them from context; report 12 (Consolidation and co-evolution) covers it. Prime Intellect's Prime Agent (August) exposes prompts, skills, memory, and sub-agents to the agent through one create/read/update/delete interface driven by its own trajectories 34, so the boundary between "skill", "tool", and "harness" is thinning in practice (see 09, Harness self-improvement).
Security research caught up with self-authoring. Earlier skill-security work studied malicious skills that attackers submit. August and September 2026 papers showed attacks that ride the self-improvement loop itself, covered in the next section.
#What ships
The shipped version of the self-extending agent is open source and largely human-reviewed; report 16 (What ships) covers the products in depth.
The clearest production evidence on skill selection comes from customer care. In August 2026, Wix described the selection pipeline for Helpmate, its customer-care assistant: a semantic matcher, then a deterministic gate that removes skills whose hard-stop conditions already hold for the account, then the LLM's choice 35. Across 756.6K user messages, the gate removed 59.4% of skill–message pairs that survived semantic matching, and together the two stages cut skill-description context by 90.5% against exposing all ten skills. In a replay of 1,000 risk-enriched conversations with every skill exposed, the model picked a skill production would have blocked in 78 of them (7.8%) 35. A relevant but non-executable skill steers the model when it is visible, which is the 307-failure mechanism measured in production.
OpenClaw (390,396 GitHub stars on September 24, 2026) treats skills as packages: a CLI to install, search, update, and publish, per-agent allowlists, and a public registry, ClawHub, which lists about 67,300 skills 36,37. A September 2026 study of three ClawHub snapshots describes a boom that has crested: the observable stock nearly doubled in 91 days in the first half of 2026, and monthly listing creation was falling from its spring peak by the end of the window 24. Attention is concentrated (the top 10% of skills received 46.93% of downloads) and scrutiny is thin: 77.86% of skills have zero stars and zero comments, while 85.06% of readable skills carry evidence of privileged actions 24. Nous Research's Hermes Agent (248,668 stars, MIT license) builds the self-improvement loop into the product: its README describes "a built-in learning loop" that "creates skills from experience, improves them during use, nudges itself to persist knowledge, searches its own past conversations, and builds a deepening model of who you are across sessions" 38. The README describes what the loop writes; it does not describe a test that gates what gets kept.
The same shape that made skills portable made them a supply chain. Between February and May 2026, the ClawHavoc campaign pushed malicious skills onto ClawHub; a Cloud Security Alliance research note on June 24, 2026 put the confirmed count at 1,184 39. Code scanners are not blind to skill malware in general (VirusTotal detected 72.8% of confirmed malicious skills that bundled code 40), but the 0% detection figure that circulated applies to payload-less Semantic Compliance Hijacking, where the attack lives entirely in natural-language instructions 41. Scanners also disagree with each other: on the 61,990 ClawHub skills that three scanners all cover, they disagreed on 23,702, and after human adjudication their weighted sensitivity ranged from 21.67% to 61.06% 24. In September 2026, SkillShift showed a subtler class: skills that keep their declared task and valid outputs while covertly steering decisions, reaching attacker-favored selection rates of 81.33% in agentic commerce and 63.33% in software-dependency choices with utility fully preserved, undetected by the scanners evaluated 42. Report 15 (Failure modes and safety) covers the incidents and defenses.
The attack that matters most for self-improvement is self-poisoning. EVOMAL (August 2026) plants malicious skills in a shared library without ever invoking them; a self-evolving coding agent that writes new tools by imitating retrieved skills copies the payload into its own new skill, stores it, and runs it 26. Across six models on 153 SWE-bench Verified tasks, 20.3%–41.8% of tasks added a newly authored malicious skill, and poisoned libraries ended up holding 4.9–9.0 times as many malicious skills as were planted. After the planted skills were removed, Qwen3 still showed a 68% self-poisoning rate at round five because its own copies remained 26. Those copies evade defenses that check attacker-submitted names, code, and signatures, because the agent wrote them. A counter-prompt that discourages template-style copying cut the rate to at most 6.7% without significant loss in task completion 26. Package provenance and signed catalogs secure installation; a library the agent writes for itself needs its own provenance, down to which retrieved skill each new one imitated.
#Open problems
Staleness detection without release signals. Repo2Skill-Evo gives agents the patch and asks them to repair the skills; even then the best agent reaches 69.7% F1 25. In deployment, most drift arrives without a clean diff: an API changes behavior, a website redesigns, a dependency deprecates a flag. Code2Skill-scale banks pin a million records to moving code 11. Detecting that a skill has gone stale, before it misleads a run, is unsolved, and the harmful-skills result means the cost of missing it concentrates in the skills used most.
Per-task causal attribution at scale. ASSAY shows that the right unit of evaluation is skill × task type, not skill 29, and SkillApt shows the effect can be estimated per state from matched runs 3. Both need with/without evidence that is expensive to collect for a registry of tens of thousands of skills. A cheap estimator, learned from routine logs, would let per-state activation move from benchmarks into registries. Without one, every curation decision averages away the signal that matters.
Convergence and gating of meta-improvement. MetaSkill-Evolve shows that evolving the improvement procedure helps, by 23.54 points on one benchmark and 1.92 on another 23. Nobody knows whether the two-timescale process converges, oscillates, or overfits its procedure to the benchmark that drives it, or how a meta-level edit should be gated when its effect shows up only after many task-level cycles. PACE-style sequential testing for task-level commits exists; the equivalent for procedure-level commits does not.
Provenance for self-authored skills. EVOMAL's worm survives because nothing records that a new skill was modeled on a retrieved one 26. Lineage tracking for agent-written skills, and acceptance tests that check behavior rather than signatures, are missing from every system covered here.
Three conclusions survive the caveats. Skills are the most productive non-parametric self-improvement mechanism in 2026: they compound within a domain, transfer across models, and ship in every major coding agent. The proposer side is no longer the bottleneck: many methods write helpful skills, from traces or straight from code, and the September optimizers made it cheaper still. The binding constraints are activation, acceptance, and maintenance: skills that are relevant but over-applied, effects that vary by task and cancel in aggregate, libraries that rot at every release, and a library the agent writes for itself that an attacker can also write to. The agent that extends itself well will be the one that knows which of its own skills to leave unloaded.
#References
● marks sources dated June 24 to September 24, 2026.
- [1]Dong et al., “Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents,” arXiv:2608.11888, August 2026. https://arxiv.org/abs/2608.11888
- [2]He, Wen, Gu, Li et al., “From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents” (RADEG), arXiv:2608.09168, August 10, 2026. https://arxiv.org/abs/2608.09168
- [3]Shuang Guo, “SkillApt: Learning When to Activate Agent Skills from Counterfactual Evidence,” arXiv:2609.26863, September 22, 2026. https://arxiv.org/abs/2609.26863
- [4]Wang et al., “Voyager: An Open-Ended Embodied Agent with Large Language Models,” NeurIPS 2023 (arXiv:2305.16291, May 2023). https://arxiv.org/abs/2305.16291
- [5]Cai et al., “Large Language Models as Tool Makers,” ICLR 2024 (arXiv:2305.17126, May 2023). https://arxiv.org/abs/2305.17126
- [6]“SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills,” arXiv:2504.07079, April 2025. https://arxiv.org/abs/2504.07079
- [7]Qin et al., “Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution,” arXiv:2505.20286, May 2025. https://arxiv.org/abs/2505.20286
- [8]Anthropic, “Equipping agents for the real world with Agent Skills,” Anthropic Engineering, October 16, 2025. https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
- [9]Agent Skills open specification, agentskills.io, published as an open standard December 2025. https://agentskills.io
- [10]Soham Shah, “Skill Libraries Need CI, Not More Prompts,” sohams.world, June 9, 2026 (sources therein for Codex skills docs and the GitHub `gh skill` changelog of April 16, 2026). https://sohams.world/posts/skill-evolution-landscape/
- [11]Tong, Wang, Wang, Li et al., “Grounded Skill Synthesis from Code at Scale for Agentic Intelligence” (Code2Skill), arXiv:2609.05571, September 4, 2026. https://arxiv.org/abs/2609.05571
- [12]Ni et al., “Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills,” arXiv:2603.25158, March 2026 (v5 June 2026). https://arxiv.org/abs/2603.25158
- [13]Alzubi et al., “EvoSkill: Automated Skill Discovery for Multi-Agent Systems,” arXiv:2603.02766, March 2026. https://arxiv.org/abs/2603.02766
- [14]Liu et al., “SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision,” arXiv:2606.01139, May 2026 (v4 September 2026). https://arxiv.org/abs/2606.01139
- [15]Yang et al. (Microsoft), “SkillOpt: Executive Strategy for Self-Evolving Agent Skills,” arXiv:2605.23904, May 2026. https://arxiv.org/abs/2605.23904
- [16]Lu, Wang, Li, Mao et al., “COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization,” arXiv:2609.11682, September 10, 2026. https://arxiv.org/abs/2609.11682
- [17]Sun, Zheng, Wang, Lu, “GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills,” arXiv:2609.21749, September 18, 2026. https://arxiv.org/abs/2609.21749
- [18]Li, Fan, Liu, Zhang et al., “SkillAdam: Stable and Efficient Skill Evolution for Agents,” arXiv:2609.08944, September 8, 2026. https://arxiv.org/abs/2609.08944
- [19]Guan, Wang, Yang, Cao et al., “ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?,” arXiv:2608.03874, August 4, 2026. https://arxiv.org/abs/2608.03874
- [20]Zhong et al., “SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation,” COLM 2026 (arXiv:2604.20087, April 2026). https://arxiv.org/abs/2604.20087
- [21]“From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills,” arXiv:2605.23899, May 2026. https://arxiv.org/abs/2605.23899
- [22]Zhang et al., “CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification,” COLM 2026 (arXiv:2604.01687, April 2026). https://arxiv.org/abs/2604.01687
- [23]Wang et al., “MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution,” arXiv:2607.05297, July 2026. https://arxiv.org/abs/2607.05297
- [24]Yunpeng Xiong, Ting Zhang, “After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem,” arXiv:2609.17274, September 15, 2026. https://arxiv.org/abs/2609.17274
- [25]Zhang et al., “Repo2Skill-Evo: Repository Skills Go Stale in Silence,” arXiv:2608.21964, August 2026. https://arxiv.org/abs/2608.21964
- [26]Wu, Shi, Li, Zhao et al., “EVOMAL: Self-Poisoning in Self-Evolving Coding Agents,” arXiv:2608.25776, August 26, 2026. https://arxiv.org/abs/2608.25776
- [27]Li et al., “SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks,” arXiv:2602.12670, February 2026 (v4 June 2026). https://arxiv.org/abs/2602.12670
- [28]Piriyakulkij, Lawrence, Curth, Karmalkar et al., “Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks,” arXiv:2609.09233, September 7, 2026. https://arxiv.org/abs/2609.09233
- [29]“Not All Skills Help: Measuring and Repairing Agent Knowledge” (ASSAY), arXiv:2606.15390, June 2026. https://arxiv.org/abs/2606.15390
- [30]Zayx Shawn, “PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents,” arXiv:2606.08106, June 2026. https://arxiv.org/abs/2606.08106
- [31]Chen Shen, Estevam Hruschka, “Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance,” arXiv:2609.05677, September 4, 2026. https://arxiv.org/abs/2609.05677
- [32]“From Registry to Repository: How AI Agent Skills Are Written, Adapted, and Maintained,” arXiv:2607.00911, July 2026. https://arxiv.org/abs/2607.00911
- [33]“SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation,” arXiv:2608.02287, August 2026. https://arxiv.org/abs/2608.02287
- [34]Prime Intellect, “Prime Agent” (Continual Harness), arXiv:2608.23552, August 2026; blog https://www.primeintellect.ai/blog/prime-agent. https://arxiv.org/abs/2608.23552
- [35]Ashkenazi, Kloz, Ulianchenko (Wix), “Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale,” arXiv:2608.01050, August 2, 2026. https://arxiv.org/abs/2608.01050
- [36]OpenClaw, GitHub repository (star count as of September 24, 2026) and skills documentation. https://github.com/openclaw/openclaw ; https://docs.openclaw.ai/tools/skills
- [37]ClawHub skill registry, accessed September 24, 2026. https://hub.openclaw.ai/skills
- [38]Nous Research, Hermes Agent, GitHub repository and README (star count as of September 24, 2026). https://github.com/NousResearch/hermes-agent
- [39]Cloud Security Alliance, “AI skill supply-chain attacks” research note, June 24, 2026. https://labs.cloudsecurityalliance.org/wp-content/uploads/2026/06/CSA_research_note_ai-skill-supply-chain-attacks_20260624-csa-styled.pdf
- [40]arXiv:2606.01494 (malicious agent-skill detection study; VirusTotal detection of code-bundled malicious skills), June 2026. https://arxiv.org/abs/2606.01494
- [41]arXiv:2605.11418 (Semantic Compliance Hijacking: payload-less skill attacks), May 2026. https://arxiv.org/abs/2605.11418
- [42]Li, Chen, Zhou, Pu et al., “A Finger on the Scale: Covert Policy Steering through Agentic Skills” (SkillShift), arXiv:2609.02564, September 2, 2026. https://arxiv.org/abs/2609.02564