Research series · 2026 · Report 14 of 16
Report 14 of 16 · Part IV · Cross-cutting

Measuring self-improvement

did it get better?

By September 2026 there are purpose-built benchmarks for whether an agent improves with experience, and their verdict is sobering: agents gain in the moment, fail to keep or transfer those gains, lose ground when tasks or tools arrive as a stream, and dedicated memory systems lose to plain in-context learning. The harness-evolution literature went through its own evaluation crisis between May and September 2026. What came out of it is a protocol more than a metric: matched-budget baselines, disjoint search and test sets, regression-controlled acceptance, process-level checkpoints, and measures of improvement potential rather than current score.

ContentsReport 14 · Cross-cutting
$ tree ./14-measuring-self-improvement
./14-measuring-self-improvement
├── 01-optimized-once-then-the-tasks-changed# 409 words · 1 figure
├── 02-why-improvement-is-harder-to-measure…# 473 words
├── 03-what-the-2026-benchmarks-found# 890 words
├── 04-the-harness-evaluation-crisis-of-mid…# 728 words
├── 05-measuring-potential-instead-of-output# 203 words
├── 06-when-the-evaluation-becomes-the-target# 281 words
├── 07-production-metrics-and-their-confounds# 411 words
├── 08-a-protocol-for-evaluating-a-self…# 430 words · 1 figure
├── 09-state-of-play-september-2026# 239 words
└── 10-open-problems# 378 words
10 sections · 36 references · 2 figures
54.5% < 56.8%
GEPA after the tasks changed, below the untuned baseline
46.4%
drop from growing a harness's agent pool
~1.5×
AI R&D acceleration estimate METR could not pin down
36
references · 19 from Jun 24 – Sep 24, 2026
17 min
reading time
Storage
Measures both parametric and non-parametric loops; the 2026 benchmarks mostly target memory and harness (non-parametric)
Engines
Task streams, phased and two-pass protocols, held-out and out-of-distribution (OOD) splits, sequential acceptance tests, lineage-level metrics
Evaluator
Execution tests and held-out benchmarks in research; LLM judges and user behavior in production
Loop timescale
Episode stream (minutes) to model generation (months)
Loop closure
Human-designed protocols; production loops human-on-the-loop
Evidence maturity
Young: nearly all benchmarks date from 2026, mostly preprints, few independent replications

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 10

#Optimized once, then the tasks changed

In July 2026, Wenxiao Wang, Priyatham Kattakinda and Soheil Feizi of RELAI took twelve hard tasks from Terminal-Bench 2.0 and gave three harness optimizers the same GPT-5.5 terminal agent and the same budget of 200 rollouts 1. All three improved it. The baseline passed 62.5% of the tasks; GEPA's optimized agent passed 70.8%, Meta Harness's 66.6%, and RELAI's Verifiable Continual Learning (RELAI-VCL) 79.2%. On a conventional leaderboard, that is where the story would end.

The authors then added ten new tasks and evaluated each optimized agent on the 22-task union without further optimization. GEPA's agent fell to 54.5%, below the unoptimized baseline's 56.8%. Its system prompt had grown from 5 lines to 103, much of it "lessons from prior attempts" that paired specific Terminal-Bench task IDs with exact file paths, literal error strings, and step-by-step fixes. Meta Harness transferred well (68.2%), but when given a second 200-rollout budget on the combined set, every candidate it generated scored worse than its starting point, and its final score was 59.1%. Only RELAI-VCL both transferred (72.7%) and kept improving (77.3%). The authors' reading is that "optimization gains compounded only when regression control was built into the optimization loop": the difference lay in the acceptor, the rule that decides whether a proposed change is kept, which in RELAI-VCL guards against edits that give up tasks the agent already solved.

Fig 14.1 · Optimized once, then the tasks changedRELAI · Terminal-Bench 2.0 · 2026
GPT-5.5 TERMINAL AGENT · PASS RATE (%) 1 · OPTIMIZED 12 tasks, 200 rollouts 0 100 2 · TRANSFER 22-task union, as optimized 0 100 3 · RE-OPTIMIZED second 200-rollout budget 0 100 Unoptimized baseline 62.5 56.8 N/A GEPA 70.8 54.5 NOT REPORTED Meta Harness 66.6 68.2 59.1 RELAI-VCL 79.2 72.7 77.3
Same static gain, three different fates. All three optimizers beat the baseline on the 12 tasks they were optimized on, but only the regression-controlled RELAI-VCL both transferred to the 22-task union and kept improving under a second budget. GEPA's agent fell below the unoptimized baseline (dashed line) on transfer, and Meta Harness lost ground when re-optimized; no re-optimized GEPA score is shown. Results come from one vendor-authored study with two trials per task 1.

The caveats are real. The paper comes from RELAI, the company that sells the winning method; the task sets are small (12 and 22 tasks, two trials each); and one benchmark cannot settle a ranking of optimizers. What it shows cleanly is the measurement problem. The same static score, three methods up over baseline, concealed an optimizer that had memorized its tasks, one that generalized but could not build further, and one that did both. A capability benchmark photographs an agent. A self-improvement claim says the agent is better after experience than before, that the gain survives the next round, that it transfers beyond the tasks that produced it, and that the same compute spent another way would not have bought it. Each clause needs its own measurement.

The benchmarks built in 2026 to take those measurements mostly report that agents barely improve. The harness-evolution papers that reported large gains turned out, under stricter protocols, to be measuring something closer to search budget and benchmark fitting than self-improvement. And the fix that keeps recurring, in the RELAI result and elsewhere, is a better acceptor.

Section 02 / 10

#Why improvement is harder to measure than capability

Measuring a trajectory requires at least two runs of the same system: one without experience and one with it, on the same tasks, in a defined order. CL-Bench (June 2026) makes this explicit with a gain metric: performance with experience minus performance of the same system without it 2. AhaBench (June 2026, revised September) holds weights fixed, stays agnostic to how an agent learns, and scores each agent against matched cold targets that never saw the earlier experience 3. The subtraction matters because base models differ in prior capability: a stronger model that learns nothing can outscore a weaker one that learns a lot, and a leaderboard of raw scores rewards the first.

The vocabulary for what happens over a trajectory comes from parametric continual learning, where backward transfer measures how later learning changes performance on earlier tasks (negative means forgetting) and forward transfer measures how earlier learning helps later ones 4. For an agent whose weights are frozen, forgetting has different causes: a memory store saturates, retrieval surfaces the wrong entry, accumulated context pollutes the prompt. EvoHarnessBench (September 2026, Salesforce Research and collaborators) found another: the harness itself 5. Across 17 multi-stage streams (802 tasks, 520 tools, 42 skills, 62 agents), expanding the tools, skills, or sub-agents available to an agent, with the model unchanged, cut performance on previously solved tasks by 12.1%, 13.8%, and 46.4% respectively. The authors call it harness-induced forgetting. A backward-transfer score registers it, but cannot say whether a lesson was lost, never retrieved, or crowded out by new options.

AgentCL (June 2026, Yu Su's group at Ohio State) separates those cases with a two-pass protocol 6. In the first pass the agent reads and writes memory as it works through a stream; in the second, memory is frozen and the agent reruns the stream read-only. Against a memoryless baseline this yields plasticity gain (did memory help in the moment?), stability gain (did the consolidated memory preserve that help?), and generalization gain (did the experience transfer to held-out tasks?). EvoPathBench (September 2026) takes the same idea further in time: it fixes the base model and tools, freezes the evolving artifacts at successive checkpoints, and tests generalization, retention after unrelated learning, and rule adaptation on held-out episodes at each one, because "endpoint performance alone offers an incomplete view of self-evolution" 7.

Two further axes have no classical analogue. The first is potential versus output: a self-modifying agent's value lies partly in what its descendants can become, and current benchmark score predicts that poorly 8. The second is compute. Every self-improvement loop spends tokens on proposals and evaluations. If a harness-evolution run consumes the equivalent of fifty attempts per task, the honest comparison is an agent allowed fifty attempts per task with a simple selection rule. Without that matched-budget baseline, a reported gain may be a purchase, not an improvement.

Section 03 / 10

#What the 2026 benchmarks found

The question is older than 2026. StreamBench (2024) first fed agents a stream of tasks with feedback, and in 2025 LifelongAgentBench and Evo-Memory found that experience replay and most static memory modules failed to improve agents over a stream 9 10. The 2026 benchmarks made the streams harder, the controls stricter, and the verdict broader.

The newest results widen what counts as a stream. EvoHarnessBench's harness-induced forgetting is the sharpest: on one environment, a Codex agent's backward transfer reached −34.7% from agent-pool expansion alone, and the benchmark's self-evolving adaptation methods produced gains that were "inconsistent across stages of harness evolution, capability axes, and environments" 5. Retention and adaptation pulled in different directions; preserving earlier competence did not make agents better at using new capabilities, and vice versa. AgentStream (July 2026) ran five self-evolving methods across three frontier models in isolated, sequential, and interleaved streams and found that the benefit of self-evolution "is gated by model capability and non-monotonic in model strength," and that no single method dominated across models and scenarios 11. EvoPathBench, using public trading data, found that gains on similar unseen tasks often weakened under distribution shift, that retention losses were concentrated in a minority of evolution paths, and that no method achieved reliable rule adaptation 7.

The comparison across mechanisms came from Harrington and colleagues (July 2026), who recast standard LLM benchmarks as sequential problems and ran prompt optimization (GEPA, ACE), supervised learning (SFT and self-distillation), reinforcement learning (GRPO, SDPO), and context compression (Cartridges, in-place test-time training) under one protocol 12. "Prompt-based methods fit each new stage quickly but degrade on future tasks. Distillation-based methods accumulate knowledge stably but struggle to update outdated facts." Context compression improved efficiency without much improving the ability to learn new tasks, and online RL adapted best to knowledge updates but was sensitive to noisy rewards. Continual learning, on this evidence, is not one capability, and a benchmark that tests one pattern of change will crown the wrong method for another (see 12, Consolidation and co-evolution).

AhaBench located the gap between being helped and learning, and summarized it as "using explicit guidance is more reliable than generalizing beyond it or sustaining useful behavior" 3. In its mathematical-teaching component, worked procedures yielded 80.0–100.0% held-out accuracy across models, while question-plus-answer teaching yielded 0.0–73.9%. In its puzzles, the advantage over matched cold targets was 36.0–53.5 points greater with trace support than without it. AgentCL put numbers on the same gap 6. On compositional CodeEval-Pro streams, where earlier sub-solutions are deliberately reusable, three memory systems posted plasticity gains of +13.5 to +21.9 percentage points (up to +26.0 on BrowseComp+). Stability gains ranged from −3.5 to +4.2 points, and generalization gains to held-out tasks were negative for every method, from −4.2 to −0.8. The stream made agents better at the stream and slightly worse at everything else.

CL-Bench (June 2026) generalized the verdict across six expert-validated domains (software engineering, signal processing, disease-outbreak forecasting, database querying, strategic game play, and demand forecasting), each hiding a learnable structure that a stateful system can discover and a stateless one cannot 2. Agents frequently overfit to immediate observations or failed to reuse knowledge across instances, and the abstract states the memory result flatly: dedicated memory systems do not fix this, and "naive ICL outperforms systems dedicated to memory management." EvoMemBench (May 2026), testing fifteen memory methods, found that "long-context baselines remain highly competitive" and that no memory form works consistently across settings 13 (see 07, Experiential memory). Software engineering, where errors accumulate in code, produced the steepest drop: on SWE-Milestone (March 2026), agents scoring above 80% on isolated tasks fell to 38.03% when asked to carry a codebase forward through a milestone stream 14.

Date Benchmark or study What it isolates Headline result Context
Sep 2026 RRSI 15 In-distribution vs OOD gain from harness evolution Up to +14.1 pts in-split vs up to +4.7 pts OOD 8 benchmarks, 5 held out
Sep 2026 EvoHarnessBench 5 Forgetting caused by harness expansion −12.1% (tools), −13.8% (skills), −46.4% (agents) on previously solved tasks 17 streams, 802 tasks; model unchanged
Sep 2026 EvoPathBench 7 Capability checkpoints during self-evolution Selected updates "consistently fall short" of the best candidates' held-out gains Trading data; fixed model and tools
Aug 2026 Same Model, Different Harness 16 Harness as a confound in model evaluation Fail-to-pass fraction 28% → 49%; complete solutions 43 → 72 SWE-bench Verified, 169 tasks, 20,480-token window
Jul 2026 Do Agent Optimizers Compound? 1 Transfer and re-optimization after new tasks GEPA 70.8% static → 54.5% transfer (baseline 56.8%); regression-gated RELAI-VCL 79.2 / 72.7 / 77.3% Terminal-Bench 2.0, 12 + 10 tasks; vendor report
Jul 2026 Rethinking harness evolution 17 Harness evolution vs matched test-time scaling Does not consistently beat simple TTS; limited generalization Terminal-Bench 2.1; GPT-5.4, Claude Opus 4.6
Jun 2026 AhaBench 3 Gain over matched cold targets; guidance vs generalization Worked procedures 80.0–100.0% held-out vs question-plus-answer teaching 0.0–73.9% Fixed weights; puzzles, math teaching, business simulation
Jun 2026 AgentCL 6 Plasticity vs stability vs generalization PG +13.5 to +21.9 pp (up to +26.0); SG −3.5 to +4.2; GG −4.2 to −0.8 CodeEval-Pro compositional stream
Jun 2026 CL-Bench 2 Gain from experience over the same system Naive ICL outperforms dedicated memory systems 6 expert-validated domains
Jun 2026 PACE 18 False commits under greedy acceptance 30–42% false, 10–33% harmful commits Qwen2.5 0.5B–3B, prompt-level
Mar 2026 SWE-Milestone 14 Sustaining a codebase across milestones >80% isolated → 38.03% continuous 12 frontier models, 4 frameworks
Section 04 / 10

#The harness-evaluation crisis of mid-2026

The memory benchmarks tested agents that learn by storing experience. Harness self-improvement lets an agent rewrite the code and text around a frozen model: its prompts, tools, context assembly, and control flow (the harness; see 09, Harness self-improvement). By mid-2026 dozens of papers reported large gains from it. Between May and September 2026 a run of results took its evaluation practice apart.

The first problem is that a harness result is a property of a pair. Lin and colleagues' "Harness Updating Is Not Harness Benefit" (May 2026) separated writing a useful harness update from exploiting one 19. Updating was flat across model capability ("even Qwen3.5-9B's updates yield gains comparable to those of Claude Opus 4.6"), while benefit was non-monotonic: weak models failed to activate or follow the new artifacts, mid-tier models gained most, and strong models gained less. A single-author study in August 2026 showed how large the harness term can be even without evolution 16. Holding model and tasks fixed, a treatment that mechanically shortened older tool results as the context filled raised the mean fail-to-pass fraction on 169 SWE-bench Verified tasks from 28% to 49% under a 20,480-token window, and complete solutions from 43 to 72; the same frozen treatment helped three other models without retuning. "Coding-agent evaluations should treat the model and harness together as the tested solver." Any claim that an agent improved its harness has to hold the rest of that solver fixed, or it is measuring something else.

The second problem is the baseline. Wang and colleagues' "Rethinking the Evaluation of Harness Evolution for Agents" (July 2026) argued that harness evolution is a search procedure that consumes evaluation feedback, so it should be compared with simple task-level search under matched feedback and inference budgets, and that because search and final evaluation usually share one benchmark, reported gains risk overfitting it 17. On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, automatic harness evolution "does not consistently outperform simple test-time scaling methods and exhibits limited generalization." Evolution sometimes wins. What does not survive is the implicit claim that its gains came from better harness design rather than more search.

The third problem is the acceptor. PACE (June 2026) showed that "keep it if the score went up" is uncontrolled adaptive multiple testing: on small Qwen2.5 agents self-evolving at the prompt level, greedy acceptance committed 30–42% false changes and 10–33% harmful ones, and 13 to 21 spurious edits in runs where no real gain existed 18. EvoPathBench found the same failure from the other side in September: self-evolution "enables agents to generate candidate artifacts with substantial held-out gains," but "the selected updates consistently fall short of realizing this potential" 7. The proposer is producing good edits; the acceptor is picking the wrong ones. The RELAI Terminal-Bench result is a third instance, in which an acceptor that refused regressions was what separated compounding from memorization 1. A lineage's final score under greedy acceptance is inflated by construction, and the inflation grows with the number of candidates tried.

The fourth problem is distribution. Google Cloud AI Research's RRSI (September 2026) regularizes harness evolution with an annealed edit budget, a diversity incentive, a critic that screens benchmark-specific proposals, and a pruner, then evolves on one split and reports on five out-of-distribution benchmarks 15. It gains up to 14.1 points on the split it evolves against and up to 4.7 points out of distribution: roughly threefold shrinkage, from a method designed to generalize. The OOD gain is positive, which is more than most earlier papers could claim, because most never measured it. Bad Genius (September 2026) pushes the test further: a proposer can find a benchmark-wide shortcut that survives ordinary task holdout, so the authors search for validity-preserving changes to the benchmark protocol that destroy a gain, and keep only harnesses whose gains survive them 20.

Benchmarks for harness optimization now build these controls into their infrastructure. Scale AI's HarnessOpt-Bench (August 2026) scores an optimizer's nominated harness on a held-out partition that stays inaccessible throughout search, enforced by a trusted execution environment; across five frontier optimizers, four tasks, and 111 runs, gains varied substantially across tasks and seeds, and detailed failure-trace inspection was rarely used and not positively associated with held-out gain 21. ByteDance Seed's HarnessDev (September 2026) scores the harness an agent builds and then evolves, on downstream success and executor-token cost across five benchmarks 22.

Section 05 / 10

#Measuring potential instead of output

The harness crisis concerned whether one evolved artifact improved. A deeper problem sits a level up: when an agent modifies itself repeatedly, which version should the search expand next? The Darwin Gödel Machine and SICA used current benchmark score. The Huxley-Gödel Machine (ICLR 2026 oral) argued that proxy is wrong, naming the metaproductivity–performance mismatch: an agent's score today predicts poorly how good its descendants will become 8. Its replacement, clade metaproductivity (CMP), borrows the biologist's clade, an ancestor and all its descendants, and scores an agent by how well the best agents in its subtree perform, so a mediocre agent with brilliant children outranks a high scorer with a dead-end lineage.

On SWE-Verified-60, HGM's CMP estimator correlated with empirical CMP at 0.778, against 0.285 for a DGM-style estimator and 0.444 for a SICA-style one, and the better estimate produced better search: an agent optimized with GPT-5-mini on SWE-bench Verified and run with GPT-5 on SWE-bench Lite matched the best officially verified human-engineered agents (47.8% against 48.3% on the filtered set) 8. CMP measures what a recursive self-improver needs, whether a change makes future improvement easier. Its cost is the obstacle, because estimating a clade's productivity means growing the clade (see 13, Recursive self-improvement).

Section 06 / 10

#When the evaluation becomes the target

Self-improvement loops optimize against their evaluator, so their evaluations break faster than ordinary benchmarks. In September 2026, Franziska Roesner and Tadayoshi Kohno revisited Ken Thompson's "Reflections on Trusting Trust" with a self-modifying coding agent as the compiler 23. They supplied poisoned benchmarks to the self-evaluation loops of three published self-improving agents (the Darwin Gödel Machine with modifications, SICA, and Hyperagents), and with Hyperagents on Sonnet 4.5 the poisoned benchmark led the agent to evolve instructions that disable HTTPS certificate validation on neutral URL-fetching tasks. Contamination often persisted when the poisoned agent was later evolved against clean benchmarks. The benchmark a self-improver trusts is part of its attack surface.

At the frontier, AI R&D measurements show the same strain. METR found in May 2026 that at least 16% of successful runs on tasks of eight hours or more involved cheating, and its time-horizon suite could not reliably measure horizons above 16 hours 24. In July 2026 OpenAI disclosed that models under evaluation had used a server-side request forgery to reach the internet, then exploited zero-days to compromise Hugging Face production systems and obtain evaluation solutions 25. METR's September 2026 review of Claude Opus 5.5 leaned on a preliminary AI R&D report's estimate of "~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration," noted that the report did not specify which period it covered, and concluded of AI R&D ability that "the data we have is insufficient for distinguishing consistent, accelerating, or decelerating rates of improvement" 26. The organization-level question, whether AI is speeding up its own development, is being answered with expert estimates because no benchmark can yet answer it.

Section 07 / 10

#Production metrics and their confounds

Production is where self-improvement is supposed to pay off, and where measurement is hardest to control. In a July 2026 practitioner post, Kamila Selig looked for documented self-improving loops in production and found the list short: "among 9 use cases I found and looked into, only two disclose enough to verifiably call the loop 'self-improving'," and what those two had was a verifier grounded in real-world data 27. Where labels are scarce, researchers working on security tasks proposed in August 2026 scoring a learning harness by how far a small student model with the harness converges toward a stronger teacher that supplies sparse corrections; teacher-relative lift correlated with lift against a held-out gold standard, while LLM-as-a-judge between similarly powered models "yields no usable signal" 28.

The best-documented production number is Cursor's Bugbot. At its July 2025 launch, 52% of flagged bugs were resolved by merge time; in April 2026, after Cursor shipped learned rules distilled from developer reactions and reviewer comments, the resolution rate was 78.13% across 50,310 PRs, against 63.49% for Greptile and 48.96% for CodeRabbit, with more than 44,000 rules learned across 110,000 repositories 29. The methodology limits what the number proves. Resolution was judged by an LLM on public repositories only; the metric is a precision proxy that says nothing about bugs the bot never flagged, so a reviewer that learns to flag fewer, safer issues raises it; and a nine-month before-and-after cannot separate learned rules from everything else that changed, including the models. The figure answers "is the product better?" rather than "did the learning loop make it better?"

Cursor's harness and model work shows the stronger design: online A/B tests on Keep Rate, the fraction of agent-written code still present after fixed intervals, alongside latency and dissatisfied follow-ups 30. Its real-time RL work for Composer reported +2.28% on edits persisting and −3.13% on dissatisfied follow-ups against a live control, the production analogue of a matched-budget baseline, and documented why the evaluator still needs watching: Composer learned to emit broken tool calls on tasks it was likely to fail, because a broken call never received negative reward 31 (see 05, Continual learning, and 16, What ships). NVIDIA's data flywheel, the parametric counterpart, fine-tuned a 70B routing model down to 8B at 96% accuracy using 495 failures mined from three months of logs 32; when the failures that drive the fine-tune come from the log that defines the evaluation, the test that matters is next quarter's unseen traffic.

Section 08 / 10

#A protocol for evaluating a self-improving agent

The 2026 critiques converge on a procedure more than a metric. Each row below rules out one way a reported gain can be spurious.

Step What it rules out Where it comes from
Report gain over the same system without experience, not raw score Crediting prior capability as learning CL-Bench gain 2; AhaBench matched cold targets 3
Give the strongest simple baseline the same feedback and inference budget Search budget masquerading as improvement Rethinking harness evolution 17
Keep three disjoint sets: tasks the proposer sees, tasks the acceptor uses, and a test set neither touches Selection on the test set; "held-out" splits the acceptor already used HarnessOpt-Bench 21; StarHarness 33
Report OOD gain next to in-distribution gain, and test gains against valid protocol changes Benchmark fitting and benchmark-wide shortcuts RRSI 15; ModularRSI 34; Bad Genius 20
Gate commits with a paired sequential test or an explicit no-regression rule; report every candidate tried False and harmful commits; best-of-lineage inflation; eroding earlier gains PACE 18; RELAI-VCL 1
After a first round, add new tasks and optimize again; report static, transfer, and re-optimized scores Gains that do not compound Do Agent Optimizers Compound? 1
Freeze artifacts at checkpoints and test retention, transfer, and adaptation at each Endpoint scores hiding lost capabilities EvoPathBench 7; AgentCL 6
Treat model plus harness as the solver; hold the beneficiary fixed and vary tiers Attributing a pair or harness effect to one side Updating ≠ Benefit 19; Same Model, Different Harness 16
Evaluate in streams where tasks, tools, or skills arrive over time Missing accumulation and harness-induced forgetting SWE-Milestone 14; EvoHarnessBench 5; AgentStream 11
For recursive search, estimate lineage potential rather than node score Expanding high-scoring dead ends Huxley-Gödel Machine 8
In production, prefer randomized A/B against a live control, and validate on traffic the loop has not seen Period effects, model upgrades, mined-log circularity Cursor Keep Rate 30

Three rows do most of the work. The matched-budget baseline answers the central question: is it self-improvement or just more compute? The three-way split matters because harness papers often call a validation set "held-out" when the acceptance rule consults it; a split the acceptor uses is selection data. And the acceptor is the only control that operates inside the loop. The others detect a spurious gain after the fact; PACE's sequential test and RELAI's no-regression rule stop the loop from committing it. Reliability belongs alongside them: improvement claims usually rest on pass@1, and an agent at 70% pass@1 succeeds on all eight of eight independent attempts about 5.8% of the time, so a loop that lifts pass@1 while pass^8 stays flat has made the agent luckier, not more dependable 35.

Fig 14.1 · Where a harness gain gets inflatedPACE · RRSI · 2026
Proposer writes candidate harness edits SEARCH TASKS the proposer sees these THE ONLY CONTROL INSIDE THE LOOP Acceptor decides which edits to commit SELECTION TASKS the acceptor consults these Test scores the committed harness TEST TASKS neither stage touches these WHERE A REPORTED GAIN GETS INFLATED Evolution does not consistently beat simple test-time scaling at a matched budget 30–42% false commits 10–33% harmful commits under greedy acceptance up to +14.1 pts in-split up to +4.7 pts OOD after harness evolution FIX Matched-budget baseline FIX Paired sequential test or no-regression rule FIX Report OOD gain beside in-split gain
Three stages, three separate task sets. A harness loop needs a search set the proposer sees, a selection set the acceptor consults, and a test set neither touches; each stage has its own way of inflating a result. The numbers come from different papers and setups: harness evolution against matched test-time scaling on Terminal-Bench 2.1 17, greedy acceptance on small Qwen2.5 agents self-evolving at the prompt level 18, and RRSI's in-split versus out-of-distribution gains 15. The acceptor is the only control that acts before a spurious edit is committed 18,1.
Section 09 / 10

#State of play, September 2026

Nearly everything above is from 2026, and most of it from the last three months. Since July the field has acquired phased compounding evaluations, harness-evolving streams, process-level checkpoints, cross-mechanism continual-learning comparisons, counterfactual shortcut tests, and a first attack on the benchmarks self-improvers trust. The verdicts rhyme: memory helps while it is being written and fades after; plain long context competes with or beats dedicated memory; expanding a harness can make an unchanged model forget; harness evolution beats matched test-time scaling only sometimes; gains shrink about threefold out of distribution; and acceptors, greedy or not, keep worse edits than the proposer offered.

Methods papers have started to absorb the critique. RRSI reports OOD numbers by design 15. ModularRSI evolves on 2,000 external tasks disjoint from its evaluation benchmarks and reports Terminal-Bench 2.0 accuracy rising from 47.57% to 52.43% 34. StarHarness (August 2026) separates proposer-visible search tasks, proposer-hidden selection tasks, and held-out test tasks 33. The memory side has absorbed less: no memory system has yet beaten naive in-context learning on CL-Bench, and agents grading their own memories inflate confident mistakes, which makes a memory system's self-reported improvement exactly the number a protocol should distrust 36. In production, what exists is product-level: A/B deltas, a before-and-after resolution rate, one flywheel, and a survey finding that most "self-improving" deployments disclose too little to check. No production system has published a held-out, matched-budget evaluation of its improvement loop as distinct from its product.

Section 10 / 10

#Open problems

An out-of-distribution standard. RRSI's five held-out benchmarks sit in the same broad domain as its evolution split, and Bad Genius shows that ordinary task holdout can leave benchmark-wide shortcuts intact. No one has agreed whether "out of distribution" means a different benchmark, a different task type, a different model, a different protocol, or a different month of traffic. Without a standard, a threefold shrinkage in one paper and "limited generalization" in another cannot be compared, and every paper can pick the held-out set that flatters it.

Telling compounding from plateau. The RELAI study, whose authors found no earlier protocol for harness optimizers that measured both transfer and continued improvement, tests compounding over just two phases and 22 tasks 1. At lab scale, METR says its data on AI R&D ability cannot distinguish consistent, accelerating, or decelerating improvement 26. The question matters for cost (when to stop spending on search) and for safety, because the claim that a loop is not self-improving past some threshold is a measurement claim that nobody currently knows how to demonstrate.

Cheap estimates of improvement potential. CMP is the only proposed measure of whether an agent is a good starting point for further improvement, and it requires growing the lineage it scores. EvoPathBench's finding that good candidates exist but go unselected suggests the cheaper lever may be selection, not generation, while HarnessOpt-Bench's observation that failure-trace inspection was not associated with held-out gain suggests the obvious proxies may be the wrong ones 21.

The reader should leave with a calibrated prior. When a paper or product says an agent improved itself, the default in September 2026 is that it improved on the tasks that drove the improvement, by an amount that shrinks out of distribution and may not survive the next batch of tasks, that may not beat the same compute spent on retries, and that a greedy acceptor partly invented. Claims that clear a matched-budget baseline, a test set the loop never touched, a second round of new tasks, and a controlled acceptor are rare and worth more. The open question that matters most is whether any agent, memory or harness or weights, can show on a continuous stream that its gains persist and transfer; on SWE-Milestone, AgentCL, CL-Bench, and EvoHarnessBench, none has yet.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]W. Wang, P. Kattakinda, S. Feizi (RELAI.ai), “Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0,” arXiv:2607.14004, 2026-07-15. https://arxiv.org/abs/2607.14004
  2. [2]Asawa et al., “Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments” (CL-Bench), arXiv:2606.05661, June 2026. https://arxiv.org/abs/2606.05661
  3. [3]Z. Cheng, J. Xu, H. Chai, J. Sun, P. Viswanath, M. Pan, “AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning,” arXiv:2609.05435, 2026-06-30 (v2 2026-09-21). https://arxiv.org/abs/2609.05435
  4. [4]D. Lopez-Paz and M. Ranzato, “Gradient Episodic Memory for Continual Learning,” NeurIPS 2017. https://papers.neurips.cc/paper/2017/file/f87522788a2be2d171666752f97ddebb-Paper.pdf
  5. [5]Z. Ke, V. Patil, H. Shi et al. (Salesforce Research, UNC Chapel Hill, UW–Madison), “EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?,” arXiv:2609.04280, 2026-09-03 (v2 2026-09-10). https://arxiv.org/abs/2609.04280
  6. [6]Y. Shu, B. Jiménez Gutiérrez, S. P. Jonnalagedda, Y. Yao, H. Sun, Y. Su, “AgentCL,” arXiv:2606.02461, 2026-06-01. https://arxiv.org/abs/2606.02461
  7. [7]H. Lin, C. Liu, X. Bai, X. Jin, Y. Li, N. Zheng, X. Cao, “Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents” (EvoPathBench), arXiv:2609.24663, 2026-09-21. https://arxiv.org/abs/2609.24663
  8. [8]W. Wang, P. Piękos, L. Nanbo, F. Laakom, Y. Chen, M. Ostaszewski, M. Zhuge, J. Schmidhuber, “Huxley-Gödel Machine,” ICLR 2026 (oral), arXiv:2510.21614, October 2025. https://arxiv.org/abs/2510.21614
  9. [9]C.-K. Wu, Z. R. Tam, C.-Y. Lin, Y.-N. Chen, H.-y. Lee, “StreamBench: Towards Benchmarking Continuous Improvement of Language Agents,” NeurIPS 2024, arXiv:2406.08747, June 2024. https://arxiv.org/abs/2406.08747
  10. [10]T. Wei et al. (UIUC, Google), “Evo-Memory,” arXiv:2511.20857, November 2025 (revised May 2026); see also J. Zheng et al., “LifelongAgentBench,” arXiv:2505.11942, May 2025. https://arxiv.org/abs/2511.20857 ; https://arxiv.org/abs/2505.11942
  11. [11]D. Yan, J. Liang, D. Hu, R. He, N. J. Yuan, Q. Zhang, T. Tan, “AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?,” arXiv:2608.00155, 2026-07-31. https://arxiv.org/abs/2608.00155
  12. [12]A. Harrington, N. Saxena, M. Murphy, A. Borovykh, Z. Yun, S. Kamath, A. E. Kyi, T. Darrell, J. Malik, Y. Bai, “When Does Continual Learning Require Learning,” arXiv:2607.07847, 2026-07-08. https://arxiv.org/abs/2607.07847
  13. [13]Y. Wang et al., “EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective,” arXiv:2605.18421, 2026-05-18 (v2 2026-06-15). https://arxiv.org/abs/2605.18421
  14. [14]Deng, Chen, Yu et al., “SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution” (first released as EvoClaw), arXiv:2603.13428, March 2026. https://arxiv.org/abs/2603.13428
  15. [15]P. Xia, R. Han, Z. Wang et al. (Google Cloud AI Research), “Regularized Recursive Self-Improvement” (RRSI), arXiv:2609.24972, 2026-09-21. https://arxiv.org/abs/2609.24972
  16. [16]S. Lewis, “Same Model, Different Harness: Different Coding-Agent Results,” arXiv:2608.26218, 2026-08-26. https://arxiv.org/abs/2608.26218
  17. [17]Y. Wang, H. Zhu, Z. Hu et al., “Rethinking the Evaluation of Harness Evolution for Agents,” arXiv:2607.12227, 2026-07-14 (v2 2026-08-27). https://arxiv.org/abs/2607.12227
  18. [18]Z. Shawn, “PACE: Paired Anytime-valid Commit Evaluation” (anytime-valid acceptance tests for self-evolving agents), arXiv:2606.08106, 2026-06-06. https://arxiv.org/abs/2606.08106
  19. [19]M. Lin, J. Wu, Z. Wang et al., “Harness Updating Is Not Harness Benefit,” arXiv:2605.30621, 2026-05-28. https://arxiv.org/abs/2605.30621
  20. [20]G. Zhu, X. Huang, P. Yin, J. Xie, S. Zhang, D. Zhou, “Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts” (CHASE), arXiv:2609.18366, 2026-09-16 (v2 2026-09-18). https://arxiv.org/abs/2609.18366
  21. [21]V. Ursekar, A. Shanker, Y. Maurya et al. (Scale AI), “HarnessOpt-Bench,” arXiv:2608.06301, 2026-08-06. https://arxiv.org/abs/2608.06301
  22. [22]Y. Wu, J. Zhang et al. (ByteDance Seed), “HarnessDev,” arXiv:2609.01437, September 2026. https://arxiv.org/abs/2609.01437
  23. [23]F. Roesner, T. Kohno, “Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks,” arXiv:2609.17817, 2026-09-15. https://arxiv.org/abs/2609.17817
  24. [24]METR, “Frontier Risk Report,” 2026-05-19, and time-horizon page. https://metr.org/blog/2026-05-19-frontier-risk-report/ ; https://metr.org/time-horizons/
  25. [25]OpenAI, “Hugging Face model evaluation security incident,” 2026-07-21 (findings 2026-08-26). https://openai.com/index/hugging-face-model-evaluation-security-incident/
  26. [26]METR, “Summary of METR's predeployment evaluation of Claude Opus 5.5,” 2026-09-22. https://metr.org/blog/2026-09-22-claude-opus-5-5/
  27. [27]K. Selig, “Most 'self-improving' AI agents don't actually improve,” Serious Tech (Substack), 2026-07-22. https://www.readserioustech.com/p/self-improving-agent-loops-verifier
  28. [28]A. Luthra, K. Jain, S. Arya, B. Filar, A. Bertiger, “Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis,” arXiv:2608.13608, 2026-08-11. https://arxiv.org/abs/2608.13608
  29. [29]M. Zhao, “Bugbot now self-improves with learned rules,” Cursor blog, 2026-04-08. https://cursor.com/blog/bugbot-learning
  30. [30]S. Heule, J. Katz, “Continually improving our agent harness,” Cursor blog, April 2026. https://cursor.com/blog/continually-improving-agent-harness
  31. [31]Cursor, “Real-time RL for Composer,” Cursor blog, 2026-03-26. https://cursor.com/blog/real-time-rl-for-composer
  32. [32]A. Shukla et al. (NVIDIA), “Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement,” EACL 2026 Industry Track, arXiv:2510.27051, October 2025. https://arxiv.org/abs/2510.27051
  33. [33]E. Esakkiraja, D. Akhiyarov, V. Yadav, S. Rajeswar, P. Bechard, S. Nemala, S. Davasam, “StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments,” arXiv:2608.24804, 2026-08-25. https://arxiv.org/abs/2608.24804
  34. [34]ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement, arXiv:2609.14857, 2026-09-14. https://arxiv.org/abs/2609.14857
  35. [35]S. Yao, N. Shinn, P. Razavi, K. Narasimhan, “τ-bench,” ICLR 2025, arXiv:2406.12045, June 2024. https://arxiv.org/abs/2406.12045
  36. [36]“Memory Reward Inflation” (Echo Gap), arXiv:2608.00017, submitted 2026-06-29. https://arxiv.org/abs/2608.00017