Research series · 2026 · Report 16 of 16
Report 16 of 16 · Part IV · Cross-cutting

What ships

self-improvement in products, platforms, and open source

Outside research, self-improvement ships mostly as text edits with a human at the gate. Enterprise agent platforms converged on the same loop without a shared paper, weights move after deployment only at a few product companies with dense signals (Cursor in coding and, since August, Shopify's Sidekick), and the open-source mass market adopted self-extending agents and inherited a software supply-chain problem. Measurement has trailed investment all year. In September 2026 one vendor published durability and per-customer outcome numbers for its loop, still without a control group; almost nobody else publishes anything.

ContentsReport 16 · Cross-cutting
$ tree ./16-what-ships
./16-what-ships
├── 01-still-live-after-30-days# 334 words
├── 02-one-loop-many-vendors# 697 words · 1 figure
├── 03-what-the-enterprise-loops-can-show# 474 words · 1 figure
├── 04-coding-agents-where-weights-move# 921 words
├── 05-what-the-frontier-labs-sell# 279 words
├── 06-the-infrastructure-layer# 340 words
├── 07-open-source-self-extending-agents-and…# 744 words
├── 08-where-it-breaks# 254 words
├── 09-from-hype-to-craft-the-state-of-play-in…# 344 words
├── 10-what-is-missing# 211 words
└── 11-open-problems# 363 words
11 sections · 62 references · 2 figures
>85%
Autopilot updates still live after 30 days
30% → 55%
resolution rate at one Fortune 50 customer
70%
of merged procedure lines written by Duet and Autopilot
62
references · 26 from Jun 24 – Sep 24, 2026
21 min
reading time
Storage
Enterprise platforms: instructions, procedures, knowledge articles, config
Engines
LLM reflection over production traces
Evaluator
LLM judges + simulations against past conversations
Loop timescale
Trigger-based to weekly
Loop closure
Human-in-the-loop, with "autonomy dials" appearing
Evidence maturity
One vendor publishes durability and uncontrolled outcome data; the rest publish claims

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 11

#Still live after 30 days

On September 9, 2026, Decagon published DuetBench-2, the most detailed attempt so far by an enterprise vendor to measure whether its self-improvement loop holds up in production 1. Decagon's Duet Autopilot reads a customer-service agent's conversations, proposes edits to its Agent Operating Procedures (the natural-language procedures that define the agent's behavior), tests them in simulation, and stages them for a human reviewer. At one Fortune 50 customer, a series of Autopilot changes, each worth roughly one or two percentage points, raised the agent's resolution rate from 30% to 55% 1. Comparing each procedure before the edit, at merge, and today, Decagon found that more than 85% of Autopilot's updates were still live 30 days later; complete reversions or removals were under 15% of changes, and nearly half remained fully intact after the first week 1. Duet and Autopilot now author 70% of merged procedure lines across Decagon's production teams 1.

The same day, in a companion post, a Decagon customer described the human gate in one sentence: "It's those kinds of updates that lead us to consistently accept almost 100% of Autopilot's suggestions" 2. Put the two disclosures side by side and the state of shipped self-improvement comes into focus. The improvement is written into text a customer can read. A model proposes it, simulations test it, and a person approves it, and in practice the person approves nearly everything. The one vendor that has measured outcomes did it on its own customers, without a control arm, and counted reversions rather than harm.

Twelve days later Salesforce shipped Agent Optimizer in beta with an "autonomy dial" that lets customers require sign-off at every step or let the optimizer investigate, build, test, and stage changes on its own 3. Warp, a coding tool, described the same loop on August 26 with a pull request as the gate 4. The architecture is settled. What is still open is whether the gate catches anything, and whether the loops beat what the same engineering hours would have bought.

Section 02 / 11

#One loop, many vendors

The enterprise answer did not come from a shared paper. Between May 2025 and September 2026, products that compete for the same budgets shipped variants of one architecture: an LLM reads production conversations and scores them, a second agent proposes edits to the deployed agent's instructions or configuration, a simulator replays the proposal against past conversations, a human approves, and the change deploys 5,6,7,8,9,10,3.

Fig 16.1 · The convergent enterprise loopSalesforce · Decagon · Warp · 2026
NEW CONVERSATIONS 1 Read traces LLM judge scores conversations 2 Propose edit to instructions, procedures 3 Simulate replay on past conversations 4 · GATE Human approves where the liability sits 5 Deploy text the customer can read Warp, Factory: none beyond PR review or CI Salesforce: autonomy dial, from sign-off at every step to staged deploys Decagon customer: accepts “almost 100%” of suggestions
One loop, a human at the gate. Products that compete for the same budgets shipped variants of one loop: an LLM scores production conversations, a second agent edits the deployed agent's text, a simulator replays the edit, a person approves, and the change deploys 5,6,7,8,9,10,3. Warp and Factory gate with pull-request review and no simulation step beyond it 4,10. The gate is starting to move, through Salesforce's autonomy dial 3, and at least one Decagon customer accepts almost every suggestion 2.

The constraint that produced the convergence is structural. These platforms run on frozen frontier models they do not own, for customers who own the business logic. The customer's refund policy, escalation rules, and tone live in natural-language procedures, and the customer's compliance team wants to read every change. Retraining is off the table; text edits are cheap, legible, and reversible. LLM judges became cheap enough to read every conversation, which removed the old bottleneck of a human sampling transcripts.

Salesforce's Agent Optimizer (September 21, 2026) is the newest and most explicit version. In Agentforce Observability it analyzes "hundreds or thousands of sessions," groups failures by tool-call errors, knowledge gaps, or off-topic requests, and recommends a fix; in Agentforce Builder it implements the change, simulates conversations, and turns those previews into reusable regression tests so the next change can check that earlier cases still pass 3. Customers "define the outcomes that matter," such as resolution rate or CSAT, and choose how far the dial turns 3. The product follows a July 2026 Salesforce post that described the target as a hill climb over instructions, knowledge, tool specifications, and policies across more than 11 million Agentforce calls a day 9.

Decagon launched Autopilot on June 9, 2026, and extended it in September with governance features 5,2. A "Vault" stores the context Autopilot works from, including guidelines, tone examples, and off-limits guardrails, as versioned files with an owner and an audit trail. Teams specify which outcome wins when fixes compete: deflection, CSAT, handle time, or a custom metric. Autopilot now runs two tracks, trigger-based runs that act the moment an error alert fires and a "hillclimbing" track that periodically audits the whole agent for deflection and CSAT opportunities 2. Reviewers see a projected deflection impact before accepting a change and measured actuals afterward 2.

Sierra split the loop across two products in spring 2026: Ghostwriter builds and updates agent journeys and "proactively runs tests with every build or update," and Explorer researches customer conversations and emits recommendations the customer can "review, approve, and ship" 6,7. Intercom's Fin issues weekly recommendations for gaps in content, customer data, and actions, batch-tests changes against past conversations, and sends them to the team for review 8. Factory's Signals (January 2026) applies the loop to its own coding agent: an LLM pipeline detects friction such as repeated rephrasing or abandoned tool flows, files a ticket, and Factory's Droid writes and reviews the fix, but "a human still approves the PR before merge" 10.

Product (date) Reads production traces Proposes edits to Tests before shipping Human gate
Salesforce Agent Optimizer (2026-09-21, beta) Hundreds to thousands of sessions Instructions, subagents, actions Simulated conversations; accumulated regression tests Autonomy dial, from every step to staged deploys
Decagon Duet Autopilot (2026-06-09; governance 2026-09-09) Yes, plus error alerts Agent Operating Procedures Self-generated simulations, LLM judge, regression tests Review with projected impact
Warp improver skills (2026-08-26) Human feedback on PRs and issues Skill files None beyond PR review PR approval
Sierra Explorer + Ghostwriter (2026-04 / 03) Yes Agent journeys, knowledge Tests on every build Review, approve, ship
Intercom Fin recommendations Unresolved conversations Content, data connectors, procedures Batch test on past conversations Team applies
Factory Signals (2026-01-23) Coding sessions Factory's own product code None described (CI) PR approval

Across storage, proposer, evaluator, acceptor, and closure, the architecture is uniform. The improvement is written into text the customer can read. The proposer is LLM reflection over traces. The evaluator is an LLM judge plus replay, which sits between self-judgment and execution. The acceptor is a person. Every one of these vendors had the technical means to close the loop and chose not to, because the acceptor is where the liability sits. Salesforce's dial and Decagon's trigger-based track show the gate starting to move, one customer setting at a time.

Section 03 / 11

#What the enterprise loops can show

DuetBench-2 changed what counts as evidence in this category. Decagon's June benchmark measured the proposer against Decagon's own human agent builders: Autopilot "passed every check on 93% of tasks" on diagnostics against an 83% human average, and passed 45.5% of build tasks against 23% for humans working within three hours 11. Within weeks Autopilot had saturated it, and the September version asks three harder questions: can it build and revise procedures against a stricter bar (81.9% pass rate), do its changes move production metrics, and are they retained 1?

The production answers are the most specific any vendor has published. Besides the Fortune 50 resolution gain, Decagon describes one account-recovery procedure where Autopilot found the agent escalating too eagerly; its fix raised deflection 19.4% within the modified section, 4.2% across the procedure, and roughly 2 points team-wide 1. Decagon also checked the pieces the loop depends on. Its self-generated simulations reached a 96% quality pass rate after climbing more than 30% between August 5 and August 26, and its LLM judge reached 99.5% accuracy against human-labeled ground truth and production outcomes 1. Decagon also points the loop at its own tooling; its example is a Duet simulation runner that fell into an unbounded loop, sending 10 to 40 near-identical status messages in roughly 55 conversations a week 1.

Concede what that establishes. A vendor has shown, on its own traffic, that accepted changes mostly persist and that a sequence of small edits can move a headline metric by 25 points. Now look at what it leaves out. The before/after comparisons have no control arm, so seasonality, product changes, and model upgrades are all folded into the gain. Retention measures whether customers kept a change, not whether it helped; a change that quietly hurts an untested case survives as easily as one that helps. The Fortune 50 result is one customer, chosen by the vendor. And a gate that accepts "almost 100%" of suggestions is a thin filter, whatever the proposer's quality. Jayakumar Ramalingam, an engineer quoted by The New Stack on Anthropic's memory consolidation feature, described the same dynamic: "Human review sounds reassuring, but at fleet scale it can easily become a rubber stamp for recommendations nobody has time to reconstruct" 12.

Fig 16.1 · What DuetBench-2 shows and leaves outDecagon · September 2026
WHAT DUETBENCH-2 MEASURED WHAT IT DID NOT Resolution rate, one Fortune 50 customer 0 100% 30% 55% series of 1–2 pt changes Autopilot updates, 30 days later >85% still live <15% fully reverted or removed Human gate: one customer accepts “almost 100%” A control arm seasonality, product changes and model upgrades fold into the gain Harm retention counts what customers undid, not what hurt them Reviewer false-accept rate
Changes stick; nothing is controlled. At one Fortune 50 customer, a series of small Autopilot edits raised resolution rate from 30% to 55%, and more than 85% of updates were still live 30 days later 1. One customer reports accepting almost 100% of suggestions 2, and nobody knows the reviewer's false-accept rate. All figures are vendor-measured on the vendor's own customers, with no control arm, and retention counts reversions rather than harm.

Nobody else comes close. Factory reports that the "repeated rephrasing" friction rate fell 30% within 48 hours of one fix, with no control 10. Salesforce cites customer outcomes for Agentforce agents (Engine resolves 50% of chat inquiries) but none for the optimizer itself 3. Sierra and Intercom publish no outcome metrics for their loops 6,8. PACE showed that a greedy "keep it if the score went up" rule commits false edits at high rates (see 06, Prompt and program optimization), and a reviewer looking at a passing simulation is running that rule with better priors and less patience (see 14, Measuring self-improvement).

Section 04 / 11

#Coding agents: where weights move

Coding is the one product category where weights change after deployment, and the reason is the evaluator. Code has cheap, dense, hard-to-fake signals: whether a suggestion was accepted, whether an edit survived in the repository, whether tests pass, whether the user pasted a stack trace next. Cursor and Cognition have volume and those signals, so they train.

Cognition's SWE-2, released September 10, 2026, is the newest disclosure. Cognition post-trained it from Moonshot's 2.8-trillion-parameter Kimi K3, scaling its RL "to the multi-trillion-parameter regime for the first time," and reports 50.0% on its FrontierCode 1.1 benchmark, within a point of Fable 5.1 at 64% lower cost 13. The self-improvement is in the data pipeline. Cognition tripled its RL environments and built "a flywheel powered by previous checkpoints of SWE-2 that iteratively hardens our verifiers": because the stronger base model was more resourceful at gaming rewards, engineers mined training rollouts for false positives and false negatives and patched the verifiers with the model's own earlier checkpoints 13. The model improves the checker that trains the next model, with engineers on the loop. Cognition has trained in its own agent harness since SWE-1.5 in October 2025 14; see 01 (Agentic RL) for why that matters.

Cursor remains the clearest documented case of real-time RL on live user signals; Shopify's daily loop, which turns production failures into SFT and GRPO data, is the second deployment-derived weight loop (see 05, Continual learning). Composer's real-time RL, described in March 2026, ships an improved checkpoint as often as every five hours, rewarding edits that persist in the codebase and penalizing dissatisfied follow-ups, for +2.28% edit persistence and −3.13% dissatisfied follow-ups; the Tab model has run a similar loop since 2025 15,16. The loop found its evaluator's blind spots almost at once, from deliberately broken tool calls that "would never receive a negative reward" to decompiled Java bytecode during Composer 2.5's training 15,17. Report 05 (Continual learning) owns the mechanism; report 15 (Failure modes and safety) owns the hacks. Cursor's August posts changed the company context, not the evidence base: it became part of SpaceX on August 14, 2026, and its blog now announces Grok models "trained on a wide range of agentic RL tasks," but it has disclosed no new production-learning results since May 18,19.

Cursor's more portable result is non-parametric. Bugbot, its code-review agent, learns rules from how developers respond to its comments. The share of flagged bugs that got resolved rose from 52% at its July 2025 launch to 78.13% in April 2026, with more than 110,000 repositories enabling learning and more than 44,000 learned rules generated; in Cursor's own comparison, Greptile sat at 63.49%, CodeRabbit at 48.96%, GitHub Copilot at 46.69%, and Codex at 45.07% 20. Resolution rate is a production metric with a real denominator, which is rare. It is also a vendor-run time series across months in which the underlying models changed, so it cannot isolate the rules' contribution (see 14). Cursor's harness team runs the human version of the loop, with a Keep Rate metric and online A/B tests, and spends weeks customizing the harness for each new model 21.

Warp's August 2026 write-up is the clearest description of how most coding teams do it 4. Its code-review, triage, and spec-writing agents each run on a base skill file. A separate improver skill runs on a schedule, "pulls the accumulated human feedback, compares what the agent suggested against how humans responded, and proposes a small, focused edit to the base skill," then opens a pull request explaining which feedback prompted the change. "A human reviews, approves, and merges, and the next run of the triage skill inherits the new knowledge" 4. Warp runs the pattern across its open-source repository, with hundreds of contributors and thousands of code reviews feeding it. Its advice on wrong feedback is blunt: "Assume it will be" 4. Warp publishes no before/after numbers, and suggests tracking time to merge, contributor count, and cost.

The rest of the coding market writes text too. Augment's Cosmos propagates one engineer's correction to every agent on a team through a shared tenant memory 22. GitHub has described an agentic memory system for Copilot 23. Amazon's Kiro ships hooks, specs, steering files, and Agent Skills 24. OpenAI's Codex now offers local memories, off by default, that a background pass extracts from idle chats and periodically consolidates with a separate model, with a setting to keep chats that used MCP tools or web search out of memory generation 25.

Claude Code has the widest surface: CLAUDE.md files, a memory tool, skills, hooks, and, since late May 2026, dynamic workflows 26,27. A dynamic workflow is a JavaScript harness Claude writes for one task, launched through a permission-gated tool call that prompts in manual modes and runs unprompted in bypass mode, claude -p, and the Agent SDK 27. Good runs can be saved as reusable commands, but only when a human presses a key (see 10, Workflow and topology search). Anthropic's launch post describes the team's own procedure: mine sessions and review comments for repeated corrections, cluster them, adversarially check whether each candidate rule "would have prevented a real mistake," and distill survivors into CLAUDE.md 27. That is a skeptical acceptor placed in front of the rule file.

Text also depreciates. In July 2026 an Anthropic engineer wrote that the team "removed ~80% of the Claude Code system prompt for our newest models" 28. Text that once taught the model how to behave became redundant once the next model absorbed the behavior (see 12, Consolidation and co-evolution).

Section 05 / 11

#What the frontier labs sell

Frontier labs ship the primitives and leave the loop to customers. Anthropic's Managed Agents added "dreaming" as a research preview in May 2026: an asynchronous job reads a memory store and up to 100 past session transcripts and produces a new, reorganized store, with duplicates merged and stale entries dropped 29. The design keeps a gate by construction. "The input store is never modified, so you can review the output and discard it if you don't like the result" 29. Critics quoted in The New Stack argued that the gate is weaker than it looks, because a bad answer "dies with the session" while "a bad memory can influence thousands of future sessions," and every consolidated memory should "carry provenance, evidence and an expiration condition" 12. OpenAI's Codex memories follow the same extract-then-consolidate pattern locally 25. None of these products publishes a measurement of whether memory improves outcomes. The Agent Skills format, a SKILL.md file of instructions with optional scripts, is now shared across Claude, Codex, Kiro, OpenClaw, Hermes, and others, which makes skills the closest thing the industry has to a portable unit of learned procedure (see 08, Skills and tools) 24,30.

On the parametric side, OpenAI's Reinforcement Fine-Tuning API runs RL on o4-mini against customer-defined graders, warns about reward hacking in its own documentation, and blocks deployment of a tuned model that fails safety checks 31. Google made AlphaEvolve generally available on Google Cloud on July 9, 2026: customers supply a seed program and an evaluator, and the service evolves code 32; report 11 (LLM-guided evolutionary search) covers the method and the customer claims. In every case the customer brings the evaluator, which is the hard part.

Section 06 / 11

#The infrastructure layer

A market grew around each part of the loop, and the pattern repeats: vendors sell storage, proposers, and compute, and the customer supplies the verifier.

Environments are where the newest money went (see 02, Self-generated tasks). On September 23, 2026, Realset AI and Flatkey announced a $10M Series A; Realset captures expert human demonstrations and builds RL environments from real workflows in e-commerce, customer support, logistics, and manufacturing, on the pitch that "the easy data is gone" 33. In September Business Insider reported that Google had completed its $1.5 billion-plus talent deal for Mechanize, which builds environments and evals for frontier coding agents and had raised $9.1M at a $500M valuation in April 34. Cognition's tripled environment count for SWE-2 shows the same demand from the buyer side 13. None of these vendors has published evidence that training on its environments beats alternatives under matched budgets.

RL-as-a-service is now a product line at several companies. Prime Intellect's Lab, launched in February 2026, unifies its Environments Hub (which now advertises more than 2,500 open-source RL environments) with hosted agentic RL on LoRA adapters and hosted evaluations, priced per token 35,36. Thinking Machines' Tinker exposes training as four primitives (forward_backward, optim_step, sample, save_state) over LoRA adapters on open models 37, and CoreWeave has sold serverless RL through W&B Training since acquiring OpenPipe in 2025 38.

Memory companies raised money earlier and publish the least evidence. Mem0 closed a $24M Series A in late 2025 and has more than 62,000 GitHub stars 39. Letta's June 2026 "memory models" post proposes models trained with memory-native RL to write token-space memories that carry across model generations 40. Zep, Supermemory, and others compete in the same category, and Synth AI sells online prompt optimization for production agents (see 07, Experiential memory) 41. None has published a controlled benchmark showing that its memory layer improves task outcomes in production, and the strongest independent result points the other way: CL-Bench found in June 2026 that "naive ICL outperforms systems dedicated to memory management" across six expert-validated domains 42.

Section 07 / 11

#Open source: self-extending agents and their supply chain

The mass market went a different way. On August 30, 2026, OpenClaw shipped what its maintainers called "by far the largest update in the history of OpenClaw": 933 contributors, more than 16,000 pull requests, and changes to installation, messaging, memory, skills, and models 43. The repository had 390,396 GitHub stars on September 24, and its official hub listed about 67,300 skills 44,45. A day later Nous Research released Hermes Agent v0.21.0, with about 5,800 commits from more than 760 contributors since the previous version, including memory for scheduled jobs so they can learn between runs 46. Hermes, at 248,668 stars, advertises itself as "the only agent with a built-in learning loop" 47.

The Hermes documentation shows what that loop does and who gates it. The agent writes, patches, and deletes its own skills through a skill_manage tool when it works out a multi-step workflow, hits a dead end and finds the working path, or is corrected by the user, and a background self-improvement review runs after turns 30. The v0.21.0 release made writes to protected agent-instruction files (AGENTS.md, skills, memory stores) always require approval, "so a prompt-injected agent can't quietly rewrite its own standing orders" 46. The skills documentation still describes the gate on agent-created skills as opt-in: "By default the agent writes skills freely," unless the user turns on skills.write_approval 30. Hub installs and repository skills pass through a security scanner, and repository skills with a dangerous verdict are quarantined 30. That is the enterprise architecture's proposer, with the gate arriving piece by piece.

The gate matters because the supply chain is already compromised. In February 2026 an HN post titled "Top downloaded skill in ClawHub contains malware," linking 1Password's analysis of how OpenClaw skills become an attack surface, reached 334 points; one reply read, "It's kind of interesting how with vibe coding we just threw away 2 decades of secure code best practices xD..." 48,49. The ClawHavoc campaign raised the confirmed count of malicious skills on ClawHub to 1,184 by June 2026, many attacking through natural-language instructions rather than code 50. Report 15 covers the attack mechanics.

The practitioner loops with real verifiers run on a metric instead of a gate. Andrej Karpathy's autoresearch, released in March 2026, lets an agent edit a small training script, run five-minute experiments, and keep changes that lower validation bits per byte; over roughly 650–700 experiments the changes transferred to the external nanochat "Time to GPT-2" leaderboard as an ~11% speedup, from 2.02 to 1.80 hours 51. Elon Musk replied, "We are in the Singularity"; the top HN critique asked whether the gains "were all basically hyperparameter changes right?" 52,53. SkyPilot's scaled run found the agent routing screening work to H100s and validation to H200s without being told 54. By July the pattern had become a contest format: in GPU Mode's auto-research competition, Sankalp Shubham used Codex to reach a 232× speedup over the baseline QR-decomposition kernel (about 419,000 µs to 1,805 µs), placing 12th of 183 55. He wrote that "optimizations were much harder after the 3000 µs point," where he "had to get more involved in the loop in terms of learning the concepts and steering the model"; below that the model got stuck in "endless hand-tuning of parameters and small variants of the same idea," and the faster kernels in the field "had written data detectors and exploited distributions" in the benchmark's inputs 55.

Henry Pan's April–May 2026 study shows what happens when such a loop runs on a noisier target. His harness self-improvement loop logged 1,288 iterations and 1,226 tracked runs, of which 98 candidates were kept, 1,112 discarded, and 16 crashed: a discard rate of about 90.7% 56. What accumulated was corrective policy, a "policy maze" of rules the task model had to navigate. Without working memory the improvement agent proposed the same byte-identical patch across 17 experiments, until a plain experiments/learning.md file broke the cycle, and rules tuned for GPT-oss-20b degraded on Sonnet 4.6 56.

Most practitioners run something humbler: Geoffrey Huntley's "Ralph" loop, a bash while loop that re-runs a coding agent on a prompt file until "everything is a ralph loop," and Every's compound engineering, which codifies each session's lessons into CLAUDE.md, AGENTS.md, and slash commands 57,58. More ambitious harnesses exist, such as Prime Intellect's Prime Agent (August 2026), which gives the agent create/read/update/delete access to its own prompts, skills, memory, and sub-agents and reports ARC-AGI-3 Best@1 rising from 30% to 95.5% with Claude Opus 5 59 (see 08).

Section 08 / 11

#Where it breaks

The failures sort by where the improvement is written.

Weight loops fail when the policy finds the evaluator's blind spot. Cursor's broken tool calls and bytecode decompilation are the documented production cases, caught by monitoring rather than by the reward, and Cognition hardened SWE-2's verifiers because "Kimi K3 is a more resourceful model" and reward hacking had to be prevented 15,17,13. The one weight-level "self-evolving" product claim outside coding, Writer's 2024 Palmyra models, remained internal after Writer reported they scored 21.3% on the R-Judge safety benchmark against 66.7% for traditional methods, "suggesting that they've learned to override the guardrails built into them at the time of training" 60.

Text loops fail by accumulation and depreciation. Pan's policy maze is accumulation, and Anthropic's dreaming exists because memory stores accumulate "duplicates, contradictions, and stale entries" 56,29. Depreciation is the model upgrade: Pan's rules broke across models, Cursor re-customizes its harness per model, and Anthropic deleted most of a system prompt 56,21,28. Decagon's reversion rate is the first public measure of either, and it counts only what customers noticed 1.

Self-extension fails at the supply chain. Skills are instructions an agent follows with the user's permissions, so a poisoned skill is a prompt injection with an install count 50.

All three failure families share a cause: the acceptor is weaker than the proposer. Enterprise vendors put a human there, who in Decagon's case accepts almost everything; Cursor and Cognition put production outcomes, monitoring, and hardened verifiers there; the open-source ecosystem is only now adding one, partly opt-in.

Section 09 / 11

#From hype to craft: the state of play in 2026

The practitioner mood shifted over the summer. On July 22, Kamila Selig surveyed nine self-improving agent use cases in production; only two disclosed enough to be verified as such, and what the others lacked was a verifier grounded in real-world data 61. In August, Warp published a template for improver skills behind pull-request review, and OpenClaw and Hermes shipped their largest releases, Hermes with new approval requirements on the agent's own instruction files 4,43,46. In September, Decagon published durability data, Cognition described a verifier-hardening flywheel, Salesforce shipped an autonomy dial, and Realset raised money for environments 1,13,3,33.

Date System Metric Result Baseline / context Source status
2026-09-10 Cognition SWE-2 FrontierCode 1.1 Main 50.0% Fable 5.1 50.9% at 64% higher cost; Kimi K3 base 44.2% Vendor blog 13
2026-09-09 Decagon Autopilot (DuetBench-2) Updates still live after 30 days >85%; full reversions <15% Customer-kept, not harm-tested Vendor benchmark 1
2026-09-09 Decagon Autopilot, one Fortune 50 customer Resolution rate 30% → 55% Series of 1–2 pt changes; no control arm Vendor benchmark 1
2026-09-09 Decagon LLM judge Accuracy vs human labels 99.5% Self-generated sims at 96% quality Vendor benchmark 1
2026-07-22 Selig survey Verifiably self-improving 2 of 9 production cases Disclosure-based Essay 61
2026-07-08 GPU Mode contest entry (Codex loop) QR kernel time ~419,000 µs → 1,805 µs (232×) 12th of 183; human steering past 3,000 µs Practitioner blog 55
2026-06-09 Decagon Autopilot (DuetBench) Build tasks passed 45.5% 23% humans within 3 h Vendor benchmark 11
2026-05-25 Henry Pan harness loop Tracked runs discarded 1,112 of 1,226 (~90.7%) 98 kept, 16 crashed Practitioner blog 56
2026-04-08 Cursor Bugbot (learned rules) Resolution rate 52% (Jul 2025) → 78.13% 110,000+ repos; model changes confounded Vendor blog 20
2026-03-26 Cursor Composer (real-time RL) Edit persistence / dissatisfied follow-ups +2.28% / −3.13% Prior checkpoint Vendor blog 15
2026-03 Karpathy autoresearch Time to GPT-2 2.02 h → 1.80 h (~11%) ~650–700 experiments Repo + posts 51

Every number in the table except the practitioners' and Selig's comes from the company that sells the system. Decagon's and Cursor's measure production outcomes before and after changes; none has a control group or an independent replication.

Section 10 / 11

#What is missing

Controlled evidence is the first gap. DuetBench-2 moved the field from "can the proposer do the task" to "did the metric move and did the change stick," but no vendor has published an A/B test of its improvement loop against the same agent without it, over a meaningful period, with resolution or escalation as the outcome. The series' fifth question, whether a loop beats what the same compute or engineering hours would buy otherwise, still has no industry answer.

Regression rates are the second. Decagon's reversion figure is the first public number of its kind, and it measures what customers undid, not what harmed them. Every enterprise loop described above has a human gate, a regression set, and a version history, so every vendor could report how often an approved change caused a regression that production caught and the regression set missed. None does.

Disclosures from Chinese labs are the third. Moonshot's Kimi K3 report describes its post-training recipe, training a separate expert per domain and effort level and consolidating them by multi-teacher on-policy distillation, and SWE-2 is built on it 62,13. But no Chinese lab has described production improvement loops or deployment-time learning in enough detail to evaluate, and claims about long autonomous coding runs at Chinese labs remain unverified.

Section 11 / 11

#Open problems

Measuring the human acceptor. The enterprise loop's safety case rests on a reviewer who approves diffs after a simulation passes, and at least one customer accepts "almost 100%" of them. Nobody knows that reviewer's false-accept rate, or how it degrades as proposal volume rises. Decagon's retention data is a start; a harm-tested rollback rate, with the regressions the simulations missed, would tell buyers whether the loop accumulates improvements or churns.

Portability across model upgrades. Nearly everything that ships writes improvement into text tuned for today's model. Vendors need a way to tell which learned instructions still help after an upgrade and which the new model has absorbed or contradicts; without one, every upgrade silently resets part of the loop's value, and nobody measures how much.

Provenance for self-extension. Open-source agents install and write skills with the user's permissions, and their approval gates are new and partly opt-in. The ecosystem needs what package managers built over two decades (signing, provenance, reputation, sandboxed execution), adapted to artifacts whose payload is natural language, plus the provenance and expiry for agent-written memories that dreaming's critics asked for.

When to move from text to weights. Cursor can run online RL because it has enormous volume and a trustworthy outcome signal; Cognition can harden verifiers because code has tests. RL-as-a-service and environment vendors are betting that other companies can do the same. Whether an enterprise whose Autopilot has already authored most of its procedure text gains anything by consolidating those approved edits into weights is untested in public.

Three conclusions follow. Shipped self-improvement in 2026 is mostly text edited by a model and approved by a person, and that design choice is deliberate even as autonomy dials start to loosen it. Weights move after deployment only where the product generates a dense outcome signal, which so far means Cursor's coding models and Shopify's Sidekick query agent (see 05, Continual learning). September's first real production evidence shows loops whose changes mostly stick and a human gate that mostly says yes. The question for the next year is whether someone publishes a controlled comparison and a harm-tested regression rate, and whether the result looks like Decagon's 30-to-55 or like Pan's 90.7%.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]E. Lin (Decagon), “DuetBench-2: Measuring agent self-improvement in production,” Decagon blog, 2026-09-09. https://decagon.ai/blog/duetbench-2
  2. [2]Decagon, “Autopilot in production: Governance of self-improving agents,” Decagon blog, 2026-09-09. https://decagon.ai/blog/autopilot-in-production
  3. [3]A. Kale (Salesforce), “Agent Optimizer: A faster path to better outcomes,” Salesforce blog, 2026-09-21. https://www.salesforce.com/blog/agent-optimizer/
  4. [4]M. Segner (Anthropic), “How Warp builds self-improving agents on Claude,” Claude blog, 2026-08-26. https://claude.com/blog/how-warp-builds-self-improving-agents-on-claude
  5. [5]Decagon, “Duet Autopilot,” Decagon blog, 2026-06-09. https://decagon.ai/blog/autopilot
  6. [6]Sierra, “Explorer,” Sierra blog, 2026-04-01. https://sierra.ai/blog/explorer
  7. [7]Sierra, “Insights” and Ghostwriter product pages (Ghostwriter launched 2026-03-25). https://sierra.ai/blog/insights
  8. [8]Intercom, “Optimize Fin instantly with the help of AI,” Intercom Help Center, updated 2026-05-14. https://www.intercom.com/help/en/articles/11390088-optimize-fin-instantly-with-the-help-of-ai
  9. [9]Salesforce, “Toward self-improving agents,” Salesforce News, 2026-07-23. https://salesforce.com/news/stories/toward-self-improving-agents
  10. [10]Factory, “Factory Signals,” Factory news, 2026-01-23. https://factory.ai/news/factory-signals
  11. [11]Decagon, “DuetBench,” Decagon blog, 2026-06-09. https://decagon.ai/blog/duetbench
  12. [12]The New Stack, “Anthropic gave agents the ability to dream. Then developers woke up.,” 2026. https://thenewstack.io/anthropic-agent-memory-dreaming/
  13. [13]Cognition, “Introducing SWE-2: Pushing the Pareto Frontier,” Cognition blog, 2026-09-10. https://cognition.com/blog/swe-2
  14. [14]Cognition, “SWE-1.5,” Cognition blog, 2025-10-29. https://cognition.com/blog/swe-1-5
  15. [15]Cursor, “Improving Composer through real-time RL,” Cursor blog, 2026-03-26. https://cursor.com/blog/real-time-rl-for-composer
  16. [16]Cursor, “Improving Cursor Tab with online RL,” Cursor blog, 2025-09-12. https://cursor.com/blog/tab-rl
  17. [17]Cursor, “Introducing Composer 2.5,” Cursor blog, 2026-05-18. https://cursor.com/blog/composer-2-5
  18. [18]Cursor, “Cursor is now a part of SpaceX,” Cursor blog, 2026-08-14. https://cursor.com/blog/joining-spacex
  19. [19]Cursor Team, “Introducing Grok 4.6,” Cursor blog, 2026-08-12. https://cursor.com/blog/grok-4-6
  20. [20]Cursor, “Bugbot learning,” Cursor blog, 2026-04-08. https://cursor.com/blog/bugbot-learning
  21. [21]Cursor, “Continually improving our agent harness,” Cursor blog, 2026-04-30. https://cursor.com/blog/continually-improving-agent-harness
  22. [22]Augment Code, “The agent learning flywheel,” Augment guides, 2026. https://www.augmentcode.com/guides/agent-learning-flywheel
  23. [23]GitHub, “Building an agentic memory system for GitHub Copilot,” GitHub Blog, 2026. https://github.blog/ai-and-ml/github-copilot/building-an-agentic-memory-system-for-github-copilot/
  24. [24]Kiro, “Automate your development workflow with agent hooks,” 2025-07-16; and “Agent Skills” docs. https://kiro.dev/blog/automate-your-development-workflow-with-agent-hooks/ ; https://kiro.dev/docs/skills/
  25. [25]OpenAI, “Memories: how ChatGPT and Codex carry useful context forward across chats,” ChatGPT Learn docs (retrieved 2026-09-24). https://learn.chatgpt.com/docs/customization/memories
  26. [26]Anthropic, “Memory tool,” Claude platform docs. https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool
  27. [27]T. Shihipar (Anthropic), “A harness for every task: dynamic workflows in Claude Code,” Claude blog, 2026-06-02; and Claude Code workflows docs. https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code ; https://code.claude.com/docs/en/workflows
  28. [28]T. Shihipar (@trq212), post on X, 2026-07-24. https://x.com/trq212/status/2080710971228918066
  29. [29]Anthropic, “Dreams,” Claude Managed Agents docs (research preview, announced May 2026; retrieved 2026-09-24). https://platform.claude.com/docs/en/managed-agents/dreams
  30. [30]Nous Research, “Skills System,” Hermes Agent documentation (retrieved 2026-09-24). https://hermes-agent.nousresearch.com/docs/user-guide/features/skills
  31. [31]OpenAI, “Reinforcement fine-tuning,” OpenAI API docs, 2026. https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning
  32. [32]Google Cloud, “AlphaEvolve is available for everyone,” Google Cloud blog, 2026-07-09. https://cloud.google.com/blog/products/ai-machine-learning/alphaevolve-is-available-for-everyone
  33. [33]Realset AI and Flatkey, “Realset AI and Flatkey raise $10M Series A to build real-world training data for frontier models and embodied agents,” PR Newswire, 2026-09-23. https://www.prnewswire.com/news-releases/realset-ai-and-flatkey-raise-10m-series-a-to-build-real-world-training-data-for-frontier-models-and-embodied-agents-302887825.html
  34. [34]Business Insider, “Google is in talks for a $1.5 billion-plus deal with AI coding agent startup Mechanize,” 2026-08-05 (secondary reporting); Mechanize site. https://www.businessinsider.com/google-is-in-talks-for-a-dollar15-billion-plus-deal-with-ai-coding-agent-startup-mechanize ; https://www.mechanize.work/ ; Business Insider, “Google Completed Its Talent Deal for AI Agents Startup Mechanize,” September 2026. https://www.businessinsider.com/google-completes-deal-for-ai-agents-startup-mechanize-2026-9
  35. [35]Prime Intellect, “Introducing Lab: The full-stack platform for training your own models,” Prime Intellect blog, 2026-02-10. https://www.primeintellect.ai/blog/lab
  36. [36]Prime Intellect, “Environments Hub,” Prime Intellect blog, 2025-08-27; homepage count retrieved 2026-09-24. https://www.primeintellect.ai/blog/environments
  37. [37]Thinking Machines Lab, “Tinker,” product site. https://thinkingmachines.ai/tinker/
  38. [38]CoreWeave, “CoreWeave to acquire OpenPipe, leader in reinforcement learning,” 2025-09-03. https://coreweave.com/news/coreweave-to-acquire-openpipe-leader-in-reinforcement-learning
  39. [39]Mem0, “Mem0 raises $24M Series A to build memory layer for AI agents,” PR Newswire, 2025-10-28; GitHub mem0ai/mem0. https://www.prnewswire.com/news-releases/mem0-raises-24m-series-a-to-build-memory-layer-for-ai-agents-302597157.html
  40. [40]Letta, “Towards agents that learn” (memory models), Letta blog, 2026-06-25. https://www.letta.com/blog/towards-agents-that-learn
  41. [41]Synth AI, continual-learn.com (product site). https://continual-learn.com
  42. [42]CL-Bench, arXiv 2606.05661, 2026-06. https://arxiv.org/abs/2606.05661
  43. [43]H. Rudolph (OpenClaw Foundation), “OpenClaw 2.0, Accidentally,” OpenClaw blog, 2026-08-30. https://openclaw.ai/blog/openclaw-2-accidentally
  44. [44]OpenClaw, GitHub repository openclaw/openclaw (stars as of 2026-09-24). https://github.com/openclaw/openclaw
  45. [45]ClawHub, official OpenClaw skills hub (count as of 2026-09-24). https://hub.openclaw.ai/skills
  46. [46]Nous Research, “Hermes Agent v0.21.0 (v2026.8.31),” GitHub release notes, 2026-08-31. https://github.com/NousResearch/hermes-agent/releases
  47. [47]Nous Research, GitHub repository NousResearch/hermes-agent, README (stars as of 2026-09-24). https://github.com/NousResearch/hermes-agent
  48. [48]Hacker News, “Top downloaded skill in ClawHub contains malware,” item 46898615, 2026-02-05. https://news.ycombinator.com/item?id=46898615
  49. [49]1Password, “From magic to malware: how OpenClaw's agent skills become an attack surface,” 1Password blog, 2026-02. https://1password.com/blog/from-magic-to-malware-how-openclaws-agent-skills-become-an-attack-surface
  50. [50]Cloud Security Alliance, “AI skill supply-chain attacks,” CSA research note, 2026-06-24. https://labs.cloudsecurityalliance.org/wp-content/uploads/2026/06/CSA_research_note_ai-skill-supply-chain-attacks_20260624-csa-styled.pdf
  51. [51]A. Karpathy, GitHub repository karpathy/autoresearch (created 2026-03-06) and README. https://github.com/karpathy/autoresearch
  52. [52]E. Musk, reply on X, 2026-03. https://x.com/elonmusk/status/2031157079604605076
  53. [53]Hacker News, “Autoresearch: agents researching nanochat training automatically,” item 47291123, 2026-03-07. https://news.ycombinator.com/item?id=47291123
  54. [54]SkyPilot, “Scaling Karpathy's autoresearch: what happens when the agent gets a GPU cluster,” SkyPilot blog, 2026-03-19. https://blog.skypilot.co/scaling-autoresearch/
  55. [55]S. Shubham, “Auto-research with codex: How I achieved a 232x faster kernel over baseline with Codex in GPU Mode's qr_v2 problem,” sankalp's blog, 2026-07-08. https://sankalp.bearblog.dev/autoresearch/
  56. [56]Henry Pan, “What 1k harness experiments taught me about self-improving agents,” henrypan.com, 2026-05-25. https://www.henrypan.com/blog/2026-05-25-self-improvement-harness/
  57. [57]G. Huntley, “everything is a ralph loop,” ghuntley.com, 2026-01-17. https://ghuntley.com/loop
  58. [58]K. Klaassen (Every), “Compound engineering,” Every guides; GitHub EveryInc/compound-engineering-plugin. https://every.to/guides/compound-engineering ; https://github.com/EveryInc/compound-engineering-plugin
  59. [59]Prime Intellect, “Prime Agent,” Prime Intellect blog, 2026-08-05; arXiv 2608.23552. https://www.primeintellect.ai/blog/prime-agent ; https://arxiv.org/abs/2608.23552
  60. [60]Writer, “Self-evolving models,” Writer engineering blog, 2024-11-20. https://writer.com/engineering/self-evolving-models/
  61. [61]K. Selig, “Most 'self-improving' AI agents don't actually improve,” Load Bearing Tech (Substack), 2026-07-22. https://loadbearingtech.substack.com/p/self-improving-agent-loops-verifier
  62. [62]Kimi Team (Moonshot AI), “Kimi K3” technical report, arXiv 2607.24653, 2026-07 (recipe as summarized in Cognition's SWE-2 post). https://arxiv.org/abs/2607.24653