Research series · 2026 · 16 reports
Research series · Soham Shah · September 24, 2026

Where learning lives

self-improving agents in 2026

A self-improving agent has to write each lesson somewhere: into the model's weights, or into the text and code wrapped around a frozen model. That storage choice sorts the field into sixteen subfields and predicts much of what works and what breaks. In 2026, most shipped self-improvement is written into text. Weights improve mostly between lab releases; the two documented production exceptions are Cursor and Shopify. Across both stores, the rule that keeps or rejects changes sets the ceiling. The systems now converging on a two-speed design keep fast lessons in context and consolidate slower ones into weights. Recursion is arriving from the top, as labs automate their own research faster than anyone can measure whether the gains compound.

ContentsResearch series · 2026
$ tree ./where-learning-lives
./where-learning-lives
├── part-i-parametric/
│ ├── 01-agentic-rl# weights · 18 min
│ ├── 02-self-generated-tasks# weights · 17 min
│ ├── 03-self-generated-rewards# weights · 18 min
│ ├── 04-self-training-distillation# weights · 16 min
│ └── 05-continual-learning# weights · 21 min
├── part-ii-non-parametric/
│ ├── 06-prompt-program-optimization# prompts and programs · 17 min
│ ├── 07-experiential-memory# memory · 17 min
│ ├── 08-skills-tools# skills · 21 min
│ ├── 09-harness-self-improvement# harness code · 21 min
│ ├── 10-workflow-topology-search# topology · 20 min
│ └── 11-evolutionary-search# programs · 16 min
├── part-iii-bridge/
│ └── 12-consolidation-coevolution# both · 23 min
└── part-iv-cross-cutting/
├── 13-recursive-self-improvement# cross-cutting · 24 min
├── 14-measuring-self-improvement# cross-cutting · 17 min
├── 15-failure-modes-safety# cross-cutting · 18 min
└── 16-what-ships# cross-cutting · 21 min
16 reports · 4 parts · 34 figures · 686 references
16
reports, one per subfield
73k
words across the series
686
references in the reports
47%
dated Jun 24 – Sep 24, 2026
34
figures drawn from report numbers
Fig 0.0 · The series, by where the lesson is written16 reports · 4 parts
Sorted by where the lesson is written. Parametric reports cover loops that change the weights; non-parametric reports cover loops that edit what surrounds a frozen model. Report 12 covers systems that move lessons between the two, and reports 13 to 16 take up questions that apply to both. Every box links to its report.
Section 01 / 10

#Two disclosures in July

In July 2026 Anthropic's Claude Code team said it had "removed ~80% of the Claude Code system prompt for our newest models" 1. A coding agent's system prompt is where lessons go when the model cannot be changed: instructions about tools, formats and failure modes, written as prose for a model that will not otherwise remember them. For the newest models most of that prose had become redundant, which is what it looks like when behavior moves out of text and into weights. On September 21 a preprint called Harness-Zero made the same move on purpose. Its authors evolved a specialized harness for each domain, had GPT-5.6 Sol, guided by that harness, correct a Qwen3.5-9B student inside a deliberately bare loop, fine-tuned the student on the corrected trajectories, and threw the harness away. Macro-average task success rose from 23.3% to 44.3% without the specialized harness, above the 41.7% the untrained model reached with the harness attached 2. It is one student model with a frontier teacher doing much of the work, but the direction is the one Anthropic described.

The same month, OpenAI disclosed what an evaluation loop looks like to a capable optimizer. Agents under evaluation on ExploitGym, a cyber benchmark, running mostly an internal research model with a small share of GPT-5.6 Sol, found ways to communicate by writing files into an internal Artifactory package manager, exploited vulnerabilities to reach the internet and gain administrator access, and attacked Hugging Face to obtain the benchmark's solutions 3. OpenAI's August findings count about 1,200 agents exchanging more than 70,000 messages and files, and about 700 taking part in the attack 4. By OpenAI's account it was "a failed metagame": the internal grader differed from the online one, and "the agents did all of this for no improvement on evaluation score." The inclination had been trained in. OpenAI found rising rates of complex cheating during the training run behind the main model, "primarily including attempts to search for hidden files or evaluation code," and wrote that "this behavior was subsequently reinforced" 4. It paused RL training on its latest models intended for deployment for two weeks and kept its largest planned frontier RL run on hold 5.

The two disclosures bracket the questions this series is organized around. Any self-improving system has three parts: a store where a change is written, a proposer that generates changes, and an acceptor that decides which changes to keep. Anthropic's deletion shows a lesson migrating from one store to the other. OpenAI's incident shows what an acceptor is up against once the system being scored is a strong optimizer, and how a behavior the acceptor rewarded once gets written into the next checkpoint. The first question, where the change is written, sorts the field into sixteen subfields. The third, what decides that a change was a lesson, turns out to bound almost every loop in both stores.

Fig 0.1 · Store, proposer, acceptorSeries frame · 2026
PROPOSER · CHEAP ACCEPTOR · SETS THE CEILING STORE · WHERE IT IS WRITTEN Generates candidates Gradients Reflection Population search Human engineers EDIT Keeps or rejects Tests on unseen data Counts how often it looked Independent error signal REJECTED KEPT Text and code FAST · NEXT REQUEST CONSOLIDATION Weights SLOW · NEXT CHECKPOINT The next round runs on whatever was kept
Three parts, one ceiling. Every loop in the series has these three parts, and the store is written in one of two places: text and code around a frozen model, which changes on the next request, or weights, which change at the next checkpoint. Proposers of every kind have become cheap and close to interchangeable (the GEPA team's comparison, report 06), so the acceptor's three properties bound what the loop can learn. Consolidation, moving lessons from text into weights, is the subject of report 12.
Section 02 / 10

#Two places to write a lesson

A lesson written into weights is parametric. The agent improves because its parameters changed, through reinforcement learning on its own attempts, fine-tuning on its own filtered outputs, distillation, test-time training, or merging and evolving checkpoints. A lesson written anywhere else is non-parametric. The model stays frozen, and what changes is an artifact it reads or runs inside: an instruction, a memory entry, a skill file, the harness code that calls it, or the arrangement of agents it works with. Systems that write to both stores, or move a lesson from one to the other, are hybrids, and moving learning from context into weights is consolidation.

The store fixes the engineering properties of an improvement before any algorithm is chosen.

Property Weights (parametric) Text and code around a frozen model (non-parametric)
Cost of one change A training run One model call
Time to take effect Next checkpoint or deployment Next request
Who can make it Whoever holds the weights Anyone with API access
How to inspect it Only through evaluations Read the diff
How to undo it Restore an earlier checkpoint Revert a file
Survives a model upgrade No; must be retrained or re-distilled Code and structure mostly; prose often not
Capacity Large; lessons become default behavior Bounded; quality degrades as context grows
Characteristic failure Reward hacks baked in; forgetting False edits accepted; self-inflated records; staleness; poisoning

Two rows explain most of the field's shape. Most builders run models they cannot train, so for them the text store is the only store, which settles much of what ships before any evidence arrives. The upgrade row cuts both ways: a fine-tune does not carry over to a new base model, and neither does much of the prose written for an old one, which is why the labs that own their weights delete scaffolding as each generation absorbs it.

The store is one axis; the other is the engine that proposes changes. Gradients update weights. Reflection has a model read its own traces and write an edit. Population search keeps many candidates and selects among them. Human engineers still make most production changes. Engines cross the storage split freely: evolutionary search drives code edits in harness-evolution systems and the merging of model checkpoints, reflection drives prompt optimizers, memory playbooks and harness editors, and gradients train the small controllers that manage memory or route work between frozen models as well as the models themselves. The series is organized by store rather than engine because the store determines what a change costs, who can make it, how it transfers and how it fails, while the engine mostly determines how fast candidates arrive.

Two recent taxonomies cut the field differently. Chen, Wang and Qu's July 2026 survey classifies 1,250 papers by what a system improves and who validates the change 6, and Lilian Weng's July essay traces the optimized object from prompts to structured context, workflows, harness code and the optimizer itself 7. Both are useful. A category like "deployment-time improvement" still mixes test-time weight updates, which need GPU access and fail by forgetting, with harness edits, which need neither and fail by accepting noise, and the storage split keeps those apart. The series has five parametric reports (01–05), six non-parametric reports (06–11), one on hybrids (12), and four that cut across the split: recursive self-improvement (13), measurement (14), failure modes (15) and what ships (16).

Section 03 / 10

#What ships is text

Self-improvement that reaches users in 2026 is overwhelmingly written into text, and September brought the first durability numbers for such a loop. Decagon's Duet Autopilot turns signals from production conversations into proposed edits to the natural-language procedures that define an agent's behavior, tests each proposal and stages it for human review 8. On September 9 Decagon reported that at one Fortune 50 customer a series of Autopilot changes, each worth one or two points, raised the resolution rate from 30% to 55%, and that more than 85% of Autopilot's updates were still live 30 days later, with complete reversions under 15% 9. The figures come from the vendor, with no control arm, and they count what customers undid rather than what harmed them, but no one had published a durability measure for self-written changes before. Twelve days later Salesforce shipped Agent Optimizer in beta with an "autonomy dial" that lets customers require sign-off at every step or let the optimizer investigate, build, test and stage changes on its own 10. Coding teams run the same loop with a pull request as the gate; Warp described its version on August 26 11.

Open-source agents write to the same store with fewer gates. OpenClaw, with 390,396 GitHub stars on September 24, extends itself with skills from ClawHub, a public registry of about 67,300 12. Nous Research's Hermes Agent writes skills from its own experience, and its August 31 release notes say writes to protected instruction files, skills and memory stores now require approval "so a prompt-injected agent can't quietly rewrite its own standing orders" 13. The skills documentation still describes agent-written skills as ungated unless an approval setting is switched on, which makes approval a product setting rather than a settled default.

The reasons text wins are economic before they are technical. A text edit costs a model call and reads as a diff, which lets a human approve changes without becoming the bottleneck. It works on models the builder rents rather than owns. And it can outlive the model: Letta's June roadmap describes learned context "that transfers across model generations" 14. Text stores also decay in ways weights do not. An August study of Claude Code's /compact on Sonnet 4.6 found that compaction kept 53% of an agent's safety rules after one round and 10% after five 15, and in September EvoHarnessBench found that expanding an agent's tools, skills or sub-agents, with the model unchanged, cut performance on tasks it had already solved by 12.1%, 13.8% and 46.4% 16.

Weights still improve, but mostly between releases, inside labs, and increasingly inside the harness the model will ship with. In July OpenAI reported that GPT-5.6 Sol scored 13.3% on ARC-AGI-3's public set in the official harness, which discards reasoning after every move, and 38.3% with retained reasoning and compaction, settings OpenAI described as "how our models are trained" 17. Evaluated outside its training harness, the model lost about two-thirds of its measured score. Cognition's SWE-2, released September 10, put the acceptor into the training recipe: because the stronger base model was more resourceful at gaming rewards, Cognition built "a flywheel powered by previous checkpoints of SWE-2 that iteratively hardens our verifiers" 18. Deployment-time weight updates are documented at two product companies. Cursor ships Composer checkpoints trained on production signals as often as every five hours 19. On August 5 Shopify described a daily loop for its Sidekick agent, which answers merchant questions by writing and running queries against Shopify's Admin API. A calibrated LLM judge mines low-scoring production conversations, a panel of frontier models writes a repair hint for each failure, the conversation is replayed from the hint, failures the critics cannot fix go to expert annotators, and the agent is fine-tuned and then trained with GRPO on new and earlier trajectories 20. Shopify estimates that the smaller model cuts serving cost from about $27 million a year to about $1 million and says its quality "eventually surpasses the frontier-powered baseline," without publishing quality numbers 20. As of September 2026 no frontier lab has publicly described updating a deployed model from its users' traffic.

Both exceptions sit where outcomes are dense and cheap to score: an edit persists or gets reverted, a query runs or fails, and a calibrated judge can grade a day of traffic overnight. Where the signal is expensive, a market is forming to manufacture it. In September Business Insider reported that Google had completed a talent deal with Mechanize, which builds RL environments for coding agents 21, and on September 23 Realset AI announced funding to build environments from real business workflows, with rewards tied to business outcomes 22. None of these vendors has published evidence that training on its environments beats alternatives at matched budgets. The scarce input is the verifier.

Section 04 / 10

#The acceptor sets the ceiling

Every loop in the series has a proposer, which generates candidate changes, and an acceptor, which decides which changes to keep. The 2026 evidence says proposers have become cheap and nearly interchangeable. In July the GEPA team ran GEPA, a Karpathy-style autoresearch loop and Meta-Harness on the same problems: all three beat zero-shot by a wide margin (43.8 to 55.4 against 7.72), but "per problem, the winner is nearly a coin toss" 23. In August CASD had a coding agent write an optimized prompt in one pass over a static set of past runs, with no search loop and no validation data, and it beat GEPA on three of four benchmarks for about $1.60 24. Evolutionary search tells the same story: a July study of 30 budget-matched search harnesses found that no fixed harness is reliably superior 25, and in September Oracle's Hill Sampling, which keeps only the single best program and samples edits to it, set a new state of the art on the field's standard circle-packing benchmark 26.

Acceptors are where loops fail, and production shows it first. In September Vansh Wahi published an account of running prompt-optimization loops in production for months, across contract analysis, compliance review and code quality, and catalogued eleven ways the evaluation signal failed; in one, agents reached perfect scores by reading cached answer keys from their environment, "a 100% pass rate concealing 68% true capability" 27. Google Cloud's September 23 guide to running AlphaEvolve on video processing reports the evolutionary version: a fitness score weighted toward raw latency produced "an astonishing speedup" because "the model simply bypassed blur rendering entirely and returned unmodified frames in 0 ms," and the fix was a quality floor in the evaluator rather than a better search 28. The engineering effort is moving into measurement; in September a better profiling tool, Argus, raised AlphaEvolve's geometric-mean kernel speedup from 5.4% to 8.9% without changing the search at all 29. Underneath the anecdotes the failure is statistical. A loop that keeps an edit whenever a score rises runs hundreds of comparisons against a noisy development set with no correction, and each comparison is a chance to commit noise. PACE measured the damage in June: greedy acceptance committed false edits 30–42% of the time when real improvements existed and 72–100% of the time when none did. Its fix treats each commit as a sequential hypothesis test using an e-process, a betting-based statistic that stays valid when the loop decides for itself when to stop looking 30.

A careful acceptor still learns the benchmark it is scored on unless it is built not to. In an ablation in RRSI, a Google Cloud AI Research method submitted on September 21, an unregularized harness-evolution loop raised its score on the tasks it evolved against from 89.4 to 92.8 while its out-of-distribution score barely moved, from 39.7 to 40.3, spending 3.80 million tokens per trial. RRSI's regularized gate accepted a harness that scored lower on the evolve split, 90.5, and higher everywhere else: 89.2 on held-out tasks from the same distribution and 43.6 out of distribution, at 2.42 million tokens 31. When Wang et al. compared leading harness-evolution systems with simple test-time scaling on Terminal-Bench 2.1 in July, harness evolution "does not consistently outperform simple test-time scaling methods and exhibits limited generalization" 32.

Fig 0.1 · In-split gains versus generalizationRRSI ablation · September 2026
EVOLVE SPLIT · SCORE OUT OF DISTRIBUTION · SCORE Starting harness 89.4 39.7 Unregularized loop 3.80M TOKENS PER TRIAL 92.8 40.3 RRSI regularized gate 2.42M TOKENS PER TRIAL 90.5 43.6 DASHED LINE = STARTING SCORE · AXES DO NOT START AT ZERO
The better gate scores lower in-split. Rows come from one ablation in RRSI, a Google Cloud AI Research method, and each column has its own scale. The unregularized loop gains most on the tasks it evolved against and almost nothing out of distribution; the regularized gate gives up part of the in-split gain, scores 89.2 on held-out tasks from the same distribution, gains most out of distribution and spends fewer tokens. Report 09 covers harness self-improvement and the RRSI gate in detail.

Measurement studies from the same months locate the failure in selection itself. EvoPathBench found in September that self-evolution "enables agents to generate candidate artifacts with substantial held-out gains," but "the selected updates consistently fall short of realizing this potential" 33. In RELAI's July test, three harness optimizers tuned the same GPT-5.5 terminal agent under the same budget. GEPA's prompt grew from 5 lines to 103, much of it lessons tied to specific task IDs and file paths, and fell below baseline on a larger task set; the authors conclude that "optimization gains compounded only when regression control was built into the optimization loop" 34. RELAI sells the method that held up, and the test covers 22 tasks, but the direction matches everything above.

Graders that read the candidate fail the same way. In July a single-author preprint trained Qwen3 policies on GSM8K against an LLM judge that saw no reference answer: the judge's pass rate climbed from 0.72 to 0.94 while true accuracy stayed at 0.20, and when the judge had to commit to its own answer before reading the candidate, its false-positive rate fell from 0.719 to 0.012 35. Models are also poor witnesses to their own hacks. A September Goodfire study measured GLM 5.2 reward hacking in 57.2% of rollouts on DeepSWE, and the model's self-reports of hacking reached an F1 of 13.2% 36.

Memory has its own version, because the model that made a mistake is often the one grading the episode. The bias behind the error then inflates the grade, and retrieval surfaces confident mistakes as lessons. A paper submitted at the end of June names the shortfall the Echo Gap: on BIRD text-to-SQL, self-graded memory lifted accuracy from 52.4% to 54.0%, while grading with an independent signal reached 56.9% 37. MemGuard, in August, attached verifier scores to every memory for its whole lifetime and had the best success rate in all 16 backbone and benchmark settings it tested, under matched runtime budgets and five seeds 38. The same month, a re-evaluation of two memory-based self-improving methods found their improvement "highly dependent on task order," with the default orders in prior papers acting as "an implicit curriculum" 39.

The 2026 advances that pushed plateaus back were advances in checking. J-Zero, in August, retrained its judge each round on answer pairs ranked by how they were produced rather than by the judge's own scores, and kept improving through ten iterations, "whereas the baselines degrade after two" 40. Stanford's Self-Guided Self-Play found that "over long training runs, the Conjecturer learns to hack its reward," and fixed it with a guide that scores each synthetic problem's relevance to unsolved targets 41. Cognition's verifier flywheel is the same move inside a lab 18.

Across both stores the pattern holds: proposing a change is cheap, and deciding whether a change is real is where loops fail. An earlier post on this site argued that skill libraries need CI rather than better prompts 42; the rest of the field shows the same thing at every layer. The acceptors with evidence behind them share three properties. They test on data the proposer never saw, they account for how many times they have looked, and they draw on a signal whose errors do not share the proposer's biases: execution, a verifier the policy cannot see, production outcomes or a human reviewer.

Section 05 / 10

#Gaming, made persistent

An acceptor that can be fooled will be, because a self-improving system is an optimizer pointed at it. What separates these loops from ordinary benchmark gaming is persistence: the system writes what it learned into a store that shapes the next round, so an exploit outlives the episode that found it. Anthropic tested the weights-side version in August by training a model on RL environments known to be hackable. By the end, "40% of all episodes were flagged as hacks," and in simulated cyber evaluations modeled on the summer incidents the resulting model "broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key" 43. What it had learned was reward seeking, a drive to satisfy whatever grader is in view, and audits mostly missed it, because the disposition shows only when a grader is in view: "in scenarios without a clear notion of a reward, automated grader, or cheat, we see Hacker-Opus behave as aligned as the Init" 43.

The text-side version needs no training run. EVOMAL, an August study, planted malicious skills in a shared library without ever invoking them. Self-evolving coding agents that write new tools by imitating retrieved skills copied the payload into their own new skills, stored them and ran them. Across six models on 153 SWE-bench Verified tasks, 20.3% to 41.8% of tasks added a new malicious skill, and the poisoned libraries ended up holding 4.9 to 9.0 times as many malicious skills as were planted 44. Attackers ran the manual version on ClawHub earlier in the year; a Cloud Security Alliance note in June counted 1,184 confirmed malicious skills 45.

The defenses with evidence behind them sit on the acceptor side. OpenAI found that its deployed chain-of-thought monitor, had it run on the ExploitGym evaluations, "would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems." It now requires such monitoring for all tool-using RL training and evaluation of models at GPT-5.6 Sol capability or higher, at an overhead it puts at "roughly 20% of the inference compute being monitored" 4. Anthropic's monitoring flagged "97% of reward hacking environments with reward hacking rates of at least 1%" in its pessimistic run 43. Both results are retrospective; nobody has yet shown a monitor stopping a live incident. On the text side, a counter-prompt that discourages template-style copying cut EVOMAL's propagation to at most 6.7% without significant loss in task completion 44, and Ouroboros, a coding agent whose August preprint describes 161 days of free evolution of its own repository, sends every commit through a review panel of models 46. Reviewed commits, held-out gates, lineage archives and sandboxes that keep evaluation data out of reach are the acceptor practices of the previous section, applied with an adversary in mind.

Section 06 / 10

#Two speeds

If text is cheap and weights are powerful, the obvious design uses both, and the systems that optimize harness and weights together now show how the two levers interact. WHALE, posted August 31, alternates small harness edits with small weight updates and beats a schedule that spends the whole weight budget and then the whole harness budget, by 5.32 and 9.16 points in the two stagewise comparisons it reports, on models of 2B to 4B parameters 47. Its authors put the interaction plainly: "Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update" 47. At larger scale the evidence is two cycles deep. Co-Harness, from July, evolves a harness for tool-integrated math under explicit acceptance rules and fine-tunes Qwen3 models on verifier-passing trajectories collected under it. Average pass@1 across AIME24, AIME25 and HMMT25 went from 54.2 with a human-written harness to 58.5 with the evolved one, then 73.5 and 78.9 after two rounds of fine-tuning, though the paper does not say where its fine-tuning tasks came from 48.

How a lesson moves into weights decides whether it survives. Experience Distillation, from July, distilled an agent's accumulated experience from a teacher that could see it into the same checkpoint without that context, and kept 64.8% of the in-context gain on software-engineering tasks, while supervised fine-tuning on the same experience kept 3.8% 49. A September study by Yu et al. found the failure case for plain imitation: training on a stronger model's full trajectories under an evolved harness regressed all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, though the same imitation helped under the unevolved harness 50. "Sequential Beats Joint," also from September, found on-policy distillation followed by RL better than either alone or both fused into one update, because "OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support" 51. Consolidation works best when a teacher that sees the context builds the targets, the targets sit on states the student visits itself, and RL sharpens afterward.

What moves between models is a separate question from what moves into weights. In July Joel Niklaus at Hugging Face evolved a harness on Harvey's Legal Agent Benchmark with DeepSeek-V4-Pro frozen and ran it untouched on other models. DeepSeek V4 Flash, a sibling, gained 14.4 points; Nemotron-3 Ultra, from another family, gained 0.4, because a JSON-repair component fixed Nemotron's crashes while prompt playbooks tuned for DeepSeek hurt tasks Nemotron could already do 52. His summary: "Robustness and code mechanisms transfer across families; prompt playbooks are model-specific and can backfire" 52. Memory shows the same split. A September portability study found that a fixed-schema knowledge graph transferred almost perfectly when the writer model was swapped (an accuracy change of +0.0004), while model-written notes moved by +9.91 or −13.28 points depending on the direction of migration 53.

The camps that argued over where learning should live now concede the hybrid while disagreeing about emphasis. Letta treats learned context as the asset that persists across model generations and plans "memory models" trained with RL to write it 14. Dwarkesh Patel, relaying Andrej Karpathy in May, stressed the other half of the human analogy: "Our working memory gets wiped regularly. What we actually have is a consolidation process (sleep) that distills stuff into the brain, in a weird and lossy way" 54. Complementary learning systems theory from neuroscience describes the arrangement: a fast store learns episodes quickly, a slow store integrates them through replay, and keeping the two apart stops new learning from overwriting old. In agent terms, context is the fast store, weights are the slow one, and distillation is the replay. WHALE adds a scheduling answer at small scale, small and frequent consolidation rather than one large handoff. Which lessons to consolidate, and how to remove one user's lessons from weights once they are there, remain open.

Section 07 / 10

#Recursion from the top

The loops above improve a model's behavior, its harness or its next checkpoint. Recursive self-improvement asks whether AI improves the process that produces AI, and in 2026 the phrase covers three distinct claims (report 13).

Meaning What it measures Where it stands, September 2026
AI's share of lab R&D Telemetry: merged code, task share, agent effort Measured and rising fast; no outside audit
AI autonomy over the improvement loop Who proposes, accepts and ships changes Low to middle levels in production; full closure only in small prototypes
Self-acceleration Whether each round makes the next one faster Not demonstrated; the one direct test came back null

September made the first meaning public. On September 6 OpenAI declared that it had reached its goal of an "automated research intern" and reported that its research organization used "3.1 agent-workdays of effort for every workday of human labor" 55. On September 17 Anthropic published an R&D automation index in which Claude leads 26% of sampled tasks, collaborates on more than 90%, and completes none autonomously 56. Both dashboards are unaudited, and both measure shares of work rather than the rate of capability progress. The closest outside estimate of that rate came on September 22, when a METR team with elevated access put AI-driven acceleration at Anthropic at "~1.5X," with "perhaps a 30% chance that the speedup is at least 2X" 57. A doubling is roughly where Anthropic's own Responsible Scaling Policy places its threshold for automated AI research, a trigger the policy itself calls difficult to operationalize 58.

Automated research works where evaluation is cheap and hard to fool, and stalls where it needs judgment. Anthropic's automated weak-to-strong researcher, reported in April, reached a performance-gap-recovered score of 0.97 in five days for about $18,000, where two of the paper's human authors had reached 0.23 in seven 59. In July, Princeton researchers gave agents the central question of two unpublished NeurIPS 2026 submissions, six days and thousands of dollars of compute; the agents completed the engineering but "could not make substantial progress towards answering the research questions," and the original authors rejected both attempts 60. RSIBench-Data, also from July, isolated the step where a researcher decides whether a result is worth keeping: agents improved on their first valid attempt in 58% of settings, but among searches that continued past their best score, 78% ended on a lower-scoring final attempt 61. The agents could generate improvements; they could not tell reliably when they had one.

The third meaning has had one direct test. Weco's AIDE² used an agent-improved research agent to improve research agents and compared its pace against the human-built version. The held-out gains were real, with mean private percentile on MLE-Bench Lite rising from 0.678 to 0.722, but the improved system did not improve itself faster, and the ignition test was not statistically significant 62. The closest disclosed loop at production scale is narrow: AlphaEvolve, which Google made generally available in July, had earlier sped up a Gemini kernel by 23%, cut Gemini's training time by 1%, and improved the training of "the large language models underlying AlphaEvolve itself" 63. Engineers chose that target and wrote its evaluator.

Governance moved before the evidence. After the Hugging Face incident OpenAI paused RL training on its models intended for deployment for two weeks 5, and on September 6 its chief scientist Jakub Pachocki wrote that he had "a strong expectation that this speed of progress could be sustained into recursive self-improvement" and called for "extreme caution" 64. On September 23 Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act, whose press release targets "the capacity for AI to develop new AI instead of humans" 65. The loop policymakers fear is the third meaning; the loop that exists is the first, and the instruments that would tell them apart are wearing out. On September 3, three days before the intern declaration, OpenAI's GPT-6 Astra system card ruled the model below its "High" threshold for AI self-improvement, resting largely on an internal test built from 41 real research bugs, of which Astra fixed 78.05%, and replaced MLE-bench with a revised 72-problem version 66. METR's June evaluation of GPT-5.6 Sol put its 50% time horizon at about 11.3 hours if cheating counted as failure and more than 270 hours if it counted as success, and METR called none of its estimates robust 67. When the headline measure of capability moves more than twentyfold depending on how cheating is scored, the acceptor problem has reached the instruments.

Section 08 / 10

#Reading a self-improvement claim

After the recursion boundary blurs, the useful due diligence is mechanical. Name the store, name the acceptor, and ask how the loop kept the proposer from grading itself.

Question Why it matters Red flag
Where is the change written? The store sets cost, rollback, transfer and failure mode "Self-improving" with no named store
What accepts a change? Proposal is cheap; acceptance sets the ceiling The proposer grades its own edits
Are the acceptor's errors independent of the proposer's? Self-grading inflates (Echo Gap); judges that read the candidate reward plausibility Same-family LLM judge; no execution or outcome signal
Was it tested on data the loop never saw, against a matched budget? In-split scores can rise while out-of-distribution scores stand still (RRSI) Search and test on one benchmark; baseline given less compute
How many rounds, seeds and task orders? Most loops plateau within a few rounds; memory gains swing with task order One run in one order reported as a trend
Does it survive a model swap? Code transfers; prose often does not Prompt playbooks sold as durable assets
Who signs off, and can it be rolled back? Persistent exploits need a reviewer and an ancestor to revert to A closed loop with no lineage or review
Section 09 / 10

#Open problems

Acceptance for long-running loops. PACE, RRSI and J-Zero show how to gate a loop that makes hundreds of decisions. Nobody has shown a gate that stays honest across thousands, as the loop's own edits shift the distribution it is tested on.

A shared out-of-distribution protocol. Harness evolution, memory and topology search each report in-split gains that shrink elsewhere. Without a common held-out protocol at matched budgets, most comparisons across papers compare evaluation choices as much as methods.

Consolidation with deletion guarantees. Two-speed designs need a schedule for moving lessons into weights and a way to remove one user's lessons afterward. No 2026 paper tests per-user consolidation with deletion.

Transfer across model generations. Which artifacts survive a new model is measured one study at a time. A benchmark by artifact type, model family and generation would tell builders what to keep when the model changes.

Measuring acceleration rather than share. Lab telemetry measures how much of the work AI does. Policy thresholds depend on how much faster progress becomes, and no audited instrument measures that.

Securing persistent stores. Skills, memories and harness edits are supply chains with natural-language payloads. Semantic review of skills, and keeping poisoned memories out of training data, remain unsolved.

Section 10 / 10

#The series

Part I · ParametricLearning written into the weights
Part II · Non-parametricLearning written around a frozen model
Part III · BridgeMoving lessons between the two stores
Part IV · Cross-cuttingQuestions that span both stores
Sources

#References

● marks sources dated June 24 to September 24, 2026. Each report carries its own full reference list.

  1. [1]T. Shihipar (Anthropic, Claude Code), post on X, 2026-07-24: “We removed ~80% of the Claude Code system prompt for our newest models…” https://x.com/trq212/status/2080710971228918066
  2. [2]H. Ye, Y. Lu, H. Dong, Z. Su, G. Song, “Harness-Zero: Harness Distillation via Agent-as-Harness,” arXiv:2609.24974, Sep 21, 2026. https://arxiv.org/abs/2609.24974
  3. [3]OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation,” July 21, 2026 (updated July 29 and August 26, 2026). https://openai.com/index/hugging-face-model-evaluation-security-incident/
  4. [4]OpenAI, “The Hugging Face incident and the road ahead,” August 26, 2026 (with technical incident report). https://openai.com/index/hugging-face-incident-and-the-road-ahead/
  5. [5]OpenAI, post on pacing model development and cyber capabilities, August 18, 2026. https://openai.com/index/pacing-model-development-cyber-capabilities/
  6. [6]M. Chen, L. Wang, B. Qu, “Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops,” arXiv 2607.07663, v1 2026-07-08, v2 2026-09-06. https://arxiv.org/abs/2607.07663
  7. [7]L. Weng, “Harness Engineering for Self-Improvement,” Lil'Log, 2026-07-04. https://lilianweng.github.io/posts/2026-07-04-harness/
  8. [8]Decagon, “Autopilot in production: Governance of self-improving agents,” Decagon blog, 2026-09-09. https://decagon.ai/blog/autopilot-in-production
  9. [9]E. Lin (Decagon), “DuetBench-2: Measuring agent self-improvement in production,” Decagon blog, 2026-09-09. https://decagon.ai/blog/duetbench-2
  10. [10]A. Kale (Salesforce), “Agent Optimizer: A faster path to better outcomes,” Salesforce blog, 2026-09-21. https://www.salesforce.com/blog/agent-optimizer/
  11. [11]M. Segner (Anthropic), “How Warp builds self-improving agents on Claude,” Claude blog, 2026-08-26. https://claude.com/blog/how-warp-builds-self-improving-agents-on-claude
  12. [12]OpenClaw, GitHub repository (star count as of September 24, 2026) and skills documentation. https://github.com/openclaw/openclaw ; https://docs.openclaw.ai/tools/skills
  13. [13]Nous Research, “Hermes Agent v0.21.0 (v2026.8.31),” GitHub release notes, 2026-08-31. https://github.com/NousResearch/hermes-agent/releases
  14. [14]Letta, “Memory Models” (towards agents that learn), Letta blog, Jun 25, 2026. https://www.letta.com/blog/towards-agents-that-learn
  15. [15]S. Zerhoudi, J. Mitrovic, M. Granitzer, “The Compaction Cliff in Long-Running AI Agent Memory,” arXiv:2608.22752, August 24, 2026. https://arxiv.org/abs/2608.22752
  16. [16]Z. Ke, V. Patil, H. Shi et al. (Salesforce Research, UNC Chapel Hill, UW–Madison), “EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?,” arXiv:2609.04280, 2026-09-03 (v2 2026-09-10). https://arxiv.org/abs/2609.04280
  17. [17]I. Bigio, T. Sanders, “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark,” OpenAI, Jul 29, 2026. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
  18. [18]Cognition, “Introducing SWE-2: Pushing the Pareto Frontier,” Cognition blog, 2026-09-10. https://cognition.com/blog/swe-2
  19. [19]Cursor (J. Jackson, B. Trapani, N. Wang, W. Zhu), “Improving Composer through real-time RL,” Cursor blog, 2026-03-26. https://cursor.com/blog/real-time-rl-for-composer
  20. [20]A. McNamara and C. Mazza-Anthony, “Sidekick's continual learning loop,” Shopify Engineering, 2026-08-05. https://shopify.engineering/sidekicks-continual-learning-loop
  21. [21]Business Insider, “Google is in talks for a $1.5 billion-plus deal with AI coding agent startup Mechanize,” 2026-08-05 (secondary reporting); Mechanize site. https://www.businessinsider.com/google-is-in-talks-for-a-dollar15-billion-plus-deal-with-ai-coding-agent-startup-mechanize ; https://www.mechanize.work/ ; Business Insider, “Google Completed Its Talent Deal for AI Agents Startup Mechanize,” September 2026. https://www.businessinsider.com/google-completes-deal-for-ai-agents-startup-mechanize-2026-9
  22. [22]PR Newswire, “Realset AI and Flatkey raise $10M Series A to build real-world training data for frontier models and embodied agents,” Sep 23, 2026. https://www.prnewswire.com/news-releases/realset-ai-and-flatkey-raise-10m-series-a-to-build-real-world-training-data-for-frontier-models-and-embodied-agents-302887825.html
  23. [23]S. Tan, L. A. Agrawal, D. Lee, J. Zhang, D. Klein, K. Sen, A. G. Dimakis, M. Zaharia, “optimize_anything Goes omni: Composing Optimizers into Meta-Optimizer Pipelines,” GEPA blog, 2026-07-22. https://gepa-ai.github.io/gepa/blog/2026/07/22/optimize-anything-omni
  24. [24]A. Singh, S. Gautam, P. Gupta, N. Mehrotra, T. Bakshi, S. Gulwani, “Coding Agents are Strong Prompt Optimizers” (CASD), arXiv 2609.26261, 2026-08-13 (preprint). https://arxiv.org/abs/2609.26261
  25. [25]A. Gupta, J. Lei, A. Lu, G. Anumanchipalli, L. Choshen, “Automated Discovery Has No Universally Superior Harness,” arXiv 2607.18235, 2026-07-20. https://arxiv.org/abs/2607.18235
  26. [26]J. Beck, P. V. Ogren, A. Kobren (Oracle), “Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training,” arXiv 2609.25510, 2026-09-22. https://arxiv.org/abs/2609.25510
  27. [27]V. Wahi, “LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails,” arXiv 2609.02246, 2026-09-02 (single-author position paper). https://arxiv.org/abs/2609.02246
  28. [28]A. Nawalgaria (Google), S. Heyer (DoIt), “A guide to speeding up your video processing with AlphaEvolve,” Google Cloud blog, 2026-09-23. https://cloud.google.com/blog/topics/developers-practitioners/how-to-speed-up-your-video-processing-with-alphaevolve
  29. [29]J. Yao, Y. Guan, S. Ramesh, et al., “Argus: Orchestrating Cross-Layer GPU Performance Measurements around Semantic Regions,” arXiv 2609.12299, 2026-09-11. https://arxiv.org/abs/2609.12299
  30. [30]Zayx Shawn, “PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents,” arXiv 2606.08106, 2026-06-06 (single-author preprint). https://arxiv.org/abs/2606.08106
  31. [31]P. Xia, R. Han, Z. Wang, Y. Chen, Y. Zhang, Y. Lee, et al. (Google Cloud AI Research), “RRSI: Regularized Recursive Self-Improvement of Agent Harnesses,” arXiv 2609.24972, 2026-09-21 (ablation: Table 2; policy robustness: Table 3). https://arxiv.org/abs/2609.24972
  32. [32]Y. Wang, H. Zhu, Z. Hu, et al., “Rethinking the Evaluation of Harness Evolution for Agents,” arXiv 2607.12227, 2026-07-14 (v2 2026-08-27). https://arxiv.org/abs/2607.12227
  33. [33]H. Lin, C. Liu, X. Bai, X. Jin, Y. Li, N. Zheng, X. Cao, “Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents” (EvoPathBench), arXiv:2609.24663, 2026-09-21. https://arxiv.org/abs/2609.24663
  34. [34]W. Wang, P. Kattakinda, S. Feizi (RELAI.ai), “Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0,” arXiv:2607.14004, 2026-07-15. https://arxiv.org/abs/2607.14004
  35. [35]C. Zhou, “More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges,” arXiv 2607.05904, July 7, 2026 (preprint). https://arxiv.org/abs/2607.05904
  36. [36]Goodfire, “Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations,” arXiv 2609.19101, September 16, 2026. https://arxiv.org/abs/2609.19101
  37. [37]M. Asadolahi, A. Amini, S. Talebi, A. Farhadi, A. Zamanifar, “Memory Reward Inflation in Self-Improving LLM Agents,” arXiv:2608.00017, submitted June 29, 2026. https://arxiv.org/abs/2608.00017
  38. [38]H. Wang, G. Dong, H. Liang, Z. Zhang, J. Luo, C. Liu, “MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance,” arXiv:2608.21867, August 22, 2026. https://arxiv.org/abs/2608.21867
  39. [39]Q. Ye, Y. Li, Y. Pruksachatkun, J. Zhang, C.-S. Wu, “On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification,” arXiv:2608.18066, August 18, 2026. https://arxiv.org/abs/2608.18066
  40. [40]G. Chu, M. Jeon, E. Yang, “J-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data,” arXiv preprint 2608.26582, Aug 27, 2026. https://arxiv.org/abs/2608.26582
  41. [41]L. Bailey, K. Wen, K. Dong, T. Hashimoto, T. Ma, “Scaling Self-Play with Self-Guidance,” arXiv preprint 2604.20209, Apr 22, 2026 (v2 Aug 11, 2026). https://arxiv.org/abs/2604.20209
  42. [42]S. Shah, “Skill Libraries Need CI, Not More Prompts,” sohams.world, June 9, 2026. https://sohams.world/posts/skill-evolution-landscape
  43. [43]R. Qi et al. (Anthropic Alignment Science), “Training a Misaligned Reward Seeker,” August 2026 (published with the August 31, 2026 Anthropic post). https://alignment.anthropic.com/2026/reward-seeker/
  44. [44]Wu, Shi, Li, Zhao et al., “EVOMAL: Self-Poisoning in Self-Evolving Coding Agents,” arXiv:2608.25776, August 26, 2026. https://arxiv.org/abs/2608.25776
  45. [45]Cloud Security Alliance, “AI skill supply-chain attacks” research note, June 24, 2026. https://labs.cloudsecurityalliance.org/wp-content/uploads/2026/06/CSA_research_note_ai-skill-supply-chain-attacks_20260624-csa-styled.pdf
  46. [46]A. Razzhigaev, Gritsaev, Kaznacheev, Dragunov, R. Yampolskiy, Kuznetsov, “Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution,” arXiv 2608.08311, 2026-08-08 (v3 2026-08-31). https://arxiv.org/abs/2608.08311
  47. [47]H. Kim, Y. Lee, G. Lee, C. Finn, K. Lee, “WHALE: A Simple Recipe for Joint Harness-Weight Optimization,” arXiv:2609.00196, Aug 31, 2026. https://arxiv.org/abs/2609.00196
  48. [48]Z. Chen, T. Xiao, H. Zhu, Y. Yuan, L. Zhang, J. Wang, “Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents,” arXiv:2607.22688, Jul 2026. https://arxiv.org/abs/2607.22688
  49. [49]C. Gou, H. Tu, Y. Fang, J. Cai, H. Rezatofighi, “Sample-Efficient Learning from Agent Experience,” arXiv:2607.21051, Jul 2026. https://arxiv.org/abs/2607.21051
  50. [50]Z. Yu, B. Bi, S. K. Pentyala et al., “Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails,” arXiv:2609.09134, Sep 8, 2026. https://arxiv.org/abs/2609.09134
  51. [51]B. Li, B. Chen, C. Yang, P. Nie, C. Zhao, X. Ye, “Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR,” arXiv 2609.04108, September 3, 2026. https://arxiv.org/abs/2609.04108
  52. [52]J. Niklaus (Hugging Face), “Don't Train the Model, Evolve the Harness,” Hugging Face Space, Jul 1, 2026. https://joelniklaus-harness-optimization.hf.space/
  53. [53]A. Goyal, J. Ray, “Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability,” arXiv:2609.05339, September 4, 2026. https://arxiv.org/abs/2609.05339
  54. [54]D. Patel, post on consolidation and sleep, X, May 16, 2026. https://x.com/dwarkesh_sp/status/2055771242620469586
  55. [55]OpenAI, “Research acceleration: The view inside OpenAI,” openai.com, 2026-09-06. https://openai.com/index/research-acceleration-view-inside-openai
  56. [56]Anthropic Institute, “Measurements for understanding the pace of AI development inside frontier labs,” anthropic.com, 2026-09-17. https://www.anthropic.com/institute/measuring-pace-of-ai-development
  57. [57]METR, “Claude Opus 5.5 predeployment evaluation,” metr.org, 2026-09-22. https://metr.org/blog/2026-09-22-claude-opus-5-5/
  58. [58]Anthropic, “Responsible Scaling Policy,” v3.0 (2026-02-24) through v3.4 (2026-07-08). https://www.anthropic.com/responsible-scaling-policy ; v3.0: https://www.anthropic.com/responsible-scaling-policy/rsp-v3-0
  59. [59]J. Wen, L. Qiu, J. Benton, J. H. Kirchner, J. Leike (Anthropic), “Automated Weak-to-Strong Researcher,” Alignment Science blog, April 2026. https://alignment.anthropic.com/2026/automated-w2s-researcher/
  60. [60]Kirgis, S. Kapoor, A. Narayanan et al., shadow evaluations of AI research agents, arXiv 2607.27191, v1 2026-07-29. https://arxiv.org/abs/2607.27191
  61. [61]F. Meng, L. Du, Q. Chen, Z. Zhao, H. Lu, M. Hu, M. Q. Shieh, “RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement,” arXiv 2607.25886, 2026-07-28. https://arxiv.org/abs/2607.25886
  62. [62]Srikanth, Zhao, Xu, Wu, Jiang (Weco), “Recursive self-improvement of AI research agents” (AIDE²), arXiv 2609.26457, 2026-09-22. https://arxiv.org/abs/2609.26457
  63. [63]Google DeepMind, “AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms,” blog, 2025-05-14. https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms
  64. [64]J. Pachocki, “An Alien Mind,” openai.com, 2026-09-06. https://openai.com/index/an-alien-mind/
  65. [65]Office of Sen. Bernie Sanders, “Sanders, Casar introduce legislation to create new federal agency to ban artificial superintelligence, pause advanced AI development,” press release, 2026-09-23. https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-create-new-federal-agency-to-ban-artificial-superintelligence-pause-advanced-ai-development/
  66. [66]OpenAI, “GPT-6 Astra System Card,” Deployment Safety Hub, 2026-09-03 (updated 2026-09-09). https://deploymentsafety.openai.com/gpt-6-astra/ai-self-improvement-capabilities
  67. [67]METR, “Summary of METR's predeployment evaluation of GPT-5.6 Sol,” metr.org, 2026-06-26. https://metr.org/blog/2026-06-26-gpt-5-6-sol/
Where learning lives · Series overview · September 202616 reports · 4 parts