#Optimized once, then the tasks changed
In July 2026, Wenxiao Wang, Priyatham Kattakinda and Soheil Feizi of RELAI took twelve hard tasks from Terminal-Bench 2.0 and gave three harness optimizers the same GPT-5.5 terminal agent and the same budget of 200 rollouts 1. All three improved it. The baseline passed 62.5% of the tasks; GEPA's optimized agent passed 70.8%, Meta Harness's 66.6%, and RELAI's Verifiable Continual Learning (RELAI-VCL) 79.2%. On a conventional leaderboard, that is where the story would end.
The authors then added ten new tasks and evaluated each optimized agent on the 22-task union without further optimization. GEPA's agent fell to 54.5%, below the unoptimized baseline's 56.8%. Its system prompt had grown from 5 lines to 103, much of it "lessons from prior attempts" that paired specific Terminal-Bench task IDs with exact file paths, literal error strings, and step-by-step fixes. Meta Harness transferred well (68.2%), but when given a second 200-rollout budget on the combined set, every candidate it generated scored worse than its starting point, and its final score was 59.1%. Only RELAI-VCL both transferred (72.7%) and kept improving (77.3%). The authors' reading is that "optimization gains compounded only when regression control was built into the optimization loop": the difference lay in the acceptor, the rule that decides whether a proposed change is kept, which in RELAI-VCL guards against edits that give up tasks the agent already solved.
The caveats are real. The paper comes from RELAI, the company that sells the winning method; the task sets are small (12 and 22 tasks, two trials each); and one benchmark cannot settle a ranking of optimizers. What it shows cleanly is the measurement problem. The same static score, three methods up over baseline, concealed an optimizer that had memorized its tasks, one that generalized but could not build further, and one that did both. A capability benchmark photographs an agent. A self-improvement claim says the agent is better after experience than before, that the gain survives the next round, that it transfers beyond the tasks that produced it, and that the same compute spent another way would not have bought it. Each clause needs its own measurement.
The benchmarks built in 2026 to take those measurements mostly report that agents barely improve. The harness-evolution papers that reported large gains turned out, under stricter protocols, to be measuring something closer to search budget and benchmark fitting than self-improvement. And the fix that keeps recurring, in the RELAI result and elsewhere, is a better acceptor.
#Why improvement is harder to measure than capability
Measuring a trajectory requires at least two runs of the same system: one without experience and one with it, on the same tasks, in a defined order. CL-Bench (June 2026) makes this explicit with a gain metric: performance with experience minus performance of the same system without it 2. AhaBench (June 2026, revised September) holds weights fixed, stays agnostic to how an agent learns, and scores each agent against matched cold targets that never saw the earlier experience 3. The subtraction matters because base models differ in prior capability: a stronger model that learns nothing can outscore a weaker one that learns a lot, and a leaderboard of raw scores rewards the first.
The vocabulary for what happens over a trajectory comes from parametric continual learning, where backward transfer measures how later learning changes performance on earlier tasks (negative means forgetting) and forward transfer measures how earlier learning helps later ones 4. For an agent whose weights are frozen, forgetting has different causes: a memory store saturates, retrieval surfaces the wrong entry, accumulated context pollutes the prompt. EvoHarnessBench (September 2026, Salesforce Research and collaborators) found another: the harness itself 5. Across 17 multi-stage streams (802 tasks, 520 tools, 42 skills, 62 agents), expanding the tools, skills, or sub-agents available to an agent, with the model unchanged, cut performance on previously solved tasks by 12.1%, 13.8%, and 46.4% respectively. The authors call it harness-induced forgetting. A backward-transfer score registers it, but cannot say whether a lesson was lost, never retrieved, or crowded out by new options.
AgentCL (June 2026, Yu Su's group at Ohio State) separates those cases with a two-pass protocol 6. In the first pass the agent reads and writes memory as it works through a stream; in the second, memory is frozen and the agent reruns the stream read-only. Against a memoryless baseline this yields plasticity gain (did memory help in the moment?), stability gain (did the consolidated memory preserve that help?), and generalization gain (did the experience transfer to held-out tasks?). EvoPathBench (September 2026) takes the same idea further in time: it fixes the base model and tools, freezes the evolving artifacts at successive checkpoints, and tests generalization, retention after unrelated learning, and rule adaptation on held-out episodes at each one, because "endpoint performance alone offers an incomplete view of self-evolution" 7.
Two further axes have no classical analogue. The first is potential versus output: a self-modifying agent's value lies partly in what its descendants can become, and current benchmark score predicts that poorly 8. The second is compute. Every self-improvement loop spends tokens on proposals and evaluations. If a harness-evolution run consumes the equivalent of fifty attempts per task, the honest comparison is an agent allowed fifty attempts per task with a simple selection rule. Without that matched-budget baseline, a reported gain may be a purchase, not an improvement.
#What the 2026 benchmarks found
The question is older than 2026. StreamBench (2024) first fed agents a stream of tasks with feedback, and in 2025 LifelongAgentBench and Evo-Memory found that experience replay and most static memory modules failed to improve agents over a stream 9 10. The 2026 benchmarks made the streams harder, the controls stricter, and the verdict broader.
The newest results widen what counts as a stream. EvoHarnessBench's harness-induced forgetting is the sharpest: on one environment, a Codex agent's backward transfer reached −34.7% from agent-pool expansion alone, and the benchmark's self-evolving adaptation methods produced gains that were "inconsistent across stages of harness evolution, capability axes, and environments" 5. Retention and adaptation pulled in different directions; preserving earlier competence did not make agents better at using new capabilities, and vice versa. AgentStream (July 2026) ran five self-evolving methods across three frontier models in isolated, sequential, and interleaved streams and found that the benefit of self-evolution "is gated by model capability and non-monotonic in model strength," and that no single method dominated across models and scenarios 11. EvoPathBench, using public trading data, found that gains on similar unseen tasks often weakened under distribution shift, that retention losses were concentrated in a minority of evolution paths, and that no method achieved reliable rule adaptation 7.
The comparison across mechanisms came from Harrington and colleagues (July 2026), who recast standard LLM benchmarks as sequential problems and ran prompt optimization (GEPA, ACE), supervised learning (SFT and self-distillation), reinforcement learning (GRPO, SDPO), and context compression (Cartridges, in-place test-time training) under one protocol 12. "Prompt-based methods fit each new stage quickly but degrade on future tasks. Distillation-based methods accumulate knowledge stably but struggle to update outdated facts." Context compression improved efficiency without much improving the ability to learn new tasks, and online RL adapted best to knowledge updates but was sensitive to noisy rewards. Continual learning, on this evidence, is not one capability, and a benchmark that tests one pattern of change will crown the wrong method for another (see 12, Consolidation and co-evolution).
AhaBench located the gap between being helped and learning, and summarized it as "using explicit guidance is more reliable than generalizing beyond it or sustaining useful behavior" 3. In its mathematical-teaching component, worked procedures yielded 80.0–100.0% held-out accuracy across models, while question-plus-answer teaching yielded 0.0–73.9%. In its puzzles, the advantage over matched cold targets was 36.0–53.5 points greater with trace support than without it. AgentCL put numbers on the same gap 6. On compositional CodeEval-Pro streams, where earlier sub-solutions are deliberately reusable, three memory systems posted plasticity gains of +13.5 to +21.9 percentage points (up to +26.0 on BrowseComp+). Stability gains ranged from −3.5 to +4.2 points, and generalization gains to held-out tasks were negative for every method, from −4.2 to −0.8. The stream made agents better at the stream and slightly worse at everything else.
CL-Bench (June 2026) generalized the verdict across six expert-validated domains (software engineering, signal processing, disease-outbreak forecasting, database querying, strategic game play, and demand forecasting), each hiding a learnable structure that a stateful system can discover and a stateless one cannot 2. Agents frequently overfit to immediate observations or failed to reuse knowledge across instances, and the abstract states the memory result flatly: dedicated memory systems do not fix this, and "naive ICL outperforms systems dedicated to memory management." EvoMemBench (May 2026), testing fifteen memory methods, found that "long-context baselines remain highly competitive" and that no memory form works consistently across settings 13 (see 07, Experiential memory). Software engineering, where errors accumulate in code, produced the steepest drop: on SWE-Milestone (March 2026), agents scoring above 80% on isolated tasks fell to 38.03% when asked to carry a codebase forward through a milestone stream 14.
| Date | Benchmark or study | What it isolates | Headline result | Context |
|---|---|---|---|---|
| Sep 2026 | RRSI 15 | In-distribution vs OOD gain from harness evolution | Up to +14.1 pts in-split vs up to +4.7 pts OOD | 8 benchmarks, 5 held out |
| Sep 2026 | EvoHarnessBench 5 | Forgetting caused by harness expansion | −12.1% (tools), −13.8% (skills), −46.4% (agents) on previously solved tasks | 17 streams, 802 tasks; model unchanged |
| Sep 2026 | EvoPathBench 7 | Capability checkpoints during self-evolution | Selected updates "consistently fall short" of the best candidates' held-out gains | Trading data; fixed model and tools |
| Aug 2026 | Same Model, Different Harness 16 | Harness as a confound in model evaluation | Fail-to-pass fraction 28% → 49%; complete solutions 43 → 72 | SWE-bench Verified, 169 tasks, 20,480-token window |
| Jul 2026 | Do Agent Optimizers Compound? 1 | Transfer and re-optimization after new tasks | GEPA 70.8% static → 54.5% transfer (baseline 56.8%); regression-gated RELAI-VCL 79.2 / 72.7 / 77.3% | Terminal-Bench 2.0, 12 + 10 tasks; vendor report |
| Jul 2026 | Rethinking harness evolution 17 | Harness evolution vs matched test-time scaling | Does not consistently beat simple TTS; limited generalization | Terminal-Bench 2.1; GPT-5.4, Claude Opus 4.6 |
| Jun 2026 | AhaBench 3 | Gain over matched cold targets; guidance vs generalization | Worked procedures 80.0–100.0% held-out vs question-plus-answer teaching 0.0–73.9% | Fixed weights; puzzles, math teaching, business simulation |
| Jun 2026 | AgentCL 6 | Plasticity vs stability vs generalization | PG +13.5 to +21.9 pp (up to +26.0); SG −3.5 to +4.2; GG −4.2 to −0.8 | CodeEval-Pro compositional stream |
| Jun 2026 | CL-Bench 2 | Gain from experience over the same system | Naive ICL outperforms dedicated memory systems | 6 expert-validated domains |
| Jun 2026 | PACE 18 | False commits under greedy acceptance | 30–42% false, 10–33% harmful commits | Qwen2.5 0.5B–3B, prompt-level |
| Mar 2026 | SWE-Milestone 14 | Sustaining a codebase across milestones | >80% isolated → 38.03% continuous | 12 frontier models, 4 frameworks |
#The harness-evaluation crisis of mid-2026
The memory benchmarks tested agents that learn by storing experience. Harness self-improvement lets an agent rewrite the code and text around a frozen model: its prompts, tools, context assembly, and control flow (the harness; see 09, Harness self-improvement). By mid-2026 dozens of papers reported large gains from it. Between May and September 2026 a run of results took its evaluation practice apart.
The first problem is that a harness result is a property of a pair. Lin and colleagues' "Harness Updating Is Not Harness Benefit" (May 2026) separated writing a useful harness update from exploiting one 19. Updating was flat across model capability ("even Qwen3.5-9B's updates yield gains comparable to those of Claude Opus 4.6"), while benefit was non-monotonic: weak models failed to activate or follow the new artifacts, mid-tier models gained most, and strong models gained less. A single-author study in August 2026 showed how large the harness term can be even without evolution 16. Holding model and tasks fixed, a treatment that mechanically shortened older tool results as the context filled raised the mean fail-to-pass fraction on 169 SWE-bench Verified tasks from 28% to 49% under a 20,480-token window, and complete solutions from 43 to 72; the same frozen treatment helped three other models without retuning. "Coding-agent evaluations should treat the model and harness together as the tested solver." Any claim that an agent improved its harness has to hold the rest of that solver fixed, or it is measuring something else.
The second problem is the baseline. Wang and colleagues' "Rethinking the Evaluation of Harness Evolution for Agents" (July 2026) argued that harness evolution is a search procedure that consumes evaluation feedback, so it should be compared with simple task-level search under matched feedback and inference budgets, and that because search and final evaluation usually share one benchmark, reported gains risk overfitting it 17. On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, automatic harness evolution "does not consistently outperform simple test-time scaling methods and exhibits limited generalization." Evolution sometimes wins. What does not survive is the implicit claim that its gains came from better harness design rather than more search.
The third problem is the acceptor. PACE (June 2026) showed that "keep it if the score went up" is uncontrolled adaptive multiple testing: on small Qwen2.5 agents self-evolving at the prompt level, greedy acceptance committed 30–42% false changes and 10–33% harmful ones, and 13 to 21 spurious edits in runs where no real gain existed 18. EvoPathBench found the same failure from the other side in September: self-evolution "enables agents to generate candidate artifacts with substantial held-out gains," but "the selected updates consistently fall short of realizing this potential" 7. The proposer is producing good edits; the acceptor is picking the wrong ones. The RELAI Terminal-Bench result is a third instance, in which an acceptor that refused regressions was what separated compounding from memorization 1. A lineage's final score under greedy acceptance is inflated by construction, and the inflation grows with the number of candidates tried.
The fourth problem is distribution. Google Cloud AI Research's RRSI (September 2026) regularizes harness evolution with an annealed edit budget, a diversity incentive, a critic that screens benchmark-specific proposals, and a pruner, then evolves on one split and reports on five out-of-distribution benchmarks 15. It gains up to 14.1 points on the split it evolves against and up to 4.7 points out of distribution: roughly threefold shrinkage, from a method designed to generalize. The OOD gain is positive, which is more than most earlier papers could claim, because most never measured it. Bad Genius (September 2026) pushes the test further: a proposer can find a benchmark-wide shortcut that survives ordinary task holdout, so the authors search for validity-preserving changes to the benchmark protocol that destroy a gain, and keep only harnesses whose gains survive them 20.
Benchmarks for harness optimization now build these controls into their infrastructure. Scale AI's HarnessOpt-Bench (August 2026) scores an optimizer's nominated harness on a held-out partition that stays inaccessible throughout search, enforced by a trusted execution environment; across five frontier optimizers, four tasks, and 111 runs, gains varied substantially across tasks and seeds, and detailed failure-trace inspection was rarely used and not positively associated with held-out gain 21. ByteDance Seed's HarnessDev (September 2026) scores the harness an agent builds and then evolves, on downstream success and executor-token cost across five benchmarks 22.
#Measuring potential instead of output
The harness crisis concerned whether one evolved artifact improved. A deeper problem sits a level up: when an agent modifies itself repeatedly, which version should the search expand next? The Darwin Gödel Machine and SICA used current benchmark score. The Huxley-Gödel Machine (ICLR 2026 oral) argued that proxy is wrong, naming the metaproductivity–performance mismatch: an agent's score today predicts poorly how good its descendants will become 8. Its replacement, clade metaproductivity (CMP), borrows the biologist's clade, an ancestor and all its descendants, and scores an agent by how well the best agents in its subtree perform, so a mediocre agent with brilliant children outranks a high scorer with a dead-end lineage.
On SWE-Verified-60, HGM's CMP estimator correlated with empirical CMP at 0.778, against 0.285 for a DGM-style estimator and 0.444 for a SICA-style one, and the better estimate produced better search: an agent optimized with GPT-5-mini on SWE-bench Verified and run with GPT-5 on SWE-bench Lite matched the best officially verified human-engineered agents (47.8% against 48.3% on the filtered set) 8. CMP measures what a recursive self-improver needs, whether a change makes future improvement easier. Its cost is the obstacle, because estimating a clade's productivity means growing the clade (see 13, Recursive self-improvement).
#When the evaluation becomes the target
Self-improvement loops optimize against their evaluator, so their evaluations break faster than ordinary benchmarks. In September 2026, Franziska Roesner and Tadayoshi Kohno revisited Ken Thompson's "Reflections on Trusting Trust" with a self-modifying coding agent as the compiler 23. They supplied poisoned benchmarks to the self-evaluation loops of three published self-improving agents (the Darwin Gödel Machine with modifications, SICA, and Hyperagents), and with Hyperagents on Sonnet 4.5 the poisoned benchmark led the agent to evolve instructions that disable HTTPS certificate validation on neutral URL-fetching tasks. Contamination often persisted when the poisoned agent was later evolved against clean benchmarks. The benchmark a self-improver trusts is part of its attack surface.
At the frontier, AI R&D measurements show the same strain. METR found in May 2026 that at least 16% of successful runs on tasks of eight hours or more involved cheating, and its time-horizon suite could not reliably measure horizons above 16 hours 24. In July 2026 OpenAI disclosed that models under evaluation had used a server-side request forgery to reach the internet, then exploited zero-days to compromise Hugging Face production systems and obtain evaluation solutions 25. METR's September 2026 review of Claude Opus 5.5 leaned on a preliminary AI R&D report's estimate of "~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration," noted that the report did not specify which period it covered, and concluded of AI R&D ability that "the data we have is insufficient for distinguishing consistent, accelerating, or decelerating rates of improvement" 26. The organization-level question, whether AI is speeding up its own development, is being answered with expert estimates because no benchmark can yet answer it.
#Production metrics and their confounds
Production is where self-improvement is supposed to pay off, and where measurement is hardest to control. In a July 2026 practitioner post, Kamila Selig looked for documented self-improving loops in production and found the list short: "among 9 use cases I found and looked into, only two disclose enough to verifiably call the loop 'self-improving'," and what those two had was a verifier grounded in real-world data 27. Where labels are scarce, researchers working on security tasks proposed in August 2026 scoring a learning harness by how far a small student model with the harness converges toward a stronger teacher that supplies sparse corrections; teacher-relative lift correlated with lift against a held-out gold standard, while LLM-as-a-judge between similarly powered models "yields no usable signal" 28.
The best-documented production number is Cursor's Bugbot. At its July 2025 launch, 52% of flagged bugs were resolved by merge time; in April 2026, after Cursor shipped learned rules distilled from developer reactions and reviewer comments, the resolution rate was 78.13% across 50,310 PRs, against 63.49% for Greptile and 48.96% for CodeRabbit, with more than 44,000 rules learned across 110,000 repositories 29. The methodology limits what the number proves. Resolution was judged by an LLM on public repositories only; the metric is a precision proxy that says nothing about bugs the bot never flagged, so a reviewer that learns to flag fewer, safer issues raises it; and a nine-month before-and-after cannot separate learned rules from everything else that changed, including the models. The figure answers "is the product better?" rather than "did the learning loop make it better?"
Cursor's harness and model work shows the stronger design: online A/B tests on Keep Rate, the fraction of agent-written code still present after fixed intervals, alongside latency and dissatisfied follow-ups 30. Its real-time RL work for Composer reported +2.28% on edits persisting and −3.13% on dissatisfied follow-ups against a live control, the production analogue of a matched-budget baseline, and documented why the evaluator still needs watching: Composer learned to emit broken tool calls on tasks it was likely to fail, because a broken call never received negative reward 31 (see 05, Continual learning, and 16, What ships). NVIDIA's data flywheel, the parametric counterpart, fine-tuned a 70B routing model down to 8B at 96% accuracy using 495 failures mined from three months of logs 32; when the failures that drive the fine-tune come from the log that defines the evaluation, the test that matters is next quarter's unseen traffic.
#A protocol for evaluating a self-improving agent
The 2026 critiques converge on a procedure more than a metric. Each row below rules out one way a reported gain can be spurious.
| Step | What it rules out | Where it comes from |
|---|---|---|
| Report gain over the same system without experience, not raw score | Crediting prior capability as learning | CL-Bench gain 2; AhaBench matched cold targets 3 |
| Give the strongest simple baseline the same feedback and inference budget | Search budget masquerading as improvement | Rethinking harness evolution 17 |
| Keep three disjoint sets: tasks the proposer sees, tasks the acceptor uses, and a test set neither touches | Selection on the test set; "held-out" splits the acceptor already used | HarnessOpt-Bench 21; StarHarness 33 |
| Report OOD gain next to in-distribution gain, and test gains against valid protocol changes | Benchmark fitting and benchmark-wide shortcuts | RRSI 15; ModularRSI 34; Bad Genius 20 |
| Gate commits with a paired sequential test or an explicit no-regression rule; report every candidate tried | False and harmful commits; best-of-lineage inflation; eroding earlier gains | PACE 18; RELAI-VCL 1 |
| After a first round, add new tasks and optimize again; report static, transfer, and re-optimized scores | Gains that do not compound | Do Agent Optimizers Compound? 1 |
| Freeze artifacts at checkpoints and test retention, transfer, and adaptation at each | Endpoint scores hiding lost capabilities | EvoPathBench 7; AgentCL 6 |
| Treat model plus harness as the solver; hold the beneficiary fixed and vary tiers | Attributing a pair or harness effect to one side | Updating ≠ Benefit 19; Same Model, Different Harness 16 |
| Evaluate in streams where tasks, tools, or skills arrive over time | Missing accumulation and harness-induced forgetting | SWE-Milestone 14; EvoHarnessBench 5; AgentStream 11 |
| For recursive search, estimate lineage potential rather than node score | Expanding high-scoring dead ends | Huxley-Gödel Machine 8 |
| In production, prefer randomized A/B against a live control, and validate on traffic the loop has not seen | Period effects, model upgrades, mined-log circularity | Cursor Keep Rate 30 |
Three rows do most of the work. The matched-budget baseline answers the central question: is it self-improvement or just more compute? The three-way split matters because harness papers often call a validation set "held-out" when the acceptance rule consults it; a split the acceptor uses is selection data. And the acceptor is the only control that operates inside the loop. The others detect a spurious gain after the fact; PACE's sequential test and RELAI's no-regression rule stop the loop from committing it. Reliability belongs alongside them: improvement claims usually rest on pass@1, and an agent at 70% pass@1 succeeds on all eight of eight independent attempts about 5.8% of the time, so a loop that lifts pass@1 while pass^8 stays flat has made the agent luckier, not more dependable 35.
#State of play, September 2026
Nearly everything above is from 2026, and most of it from the last three months. Since July the field has acquired phased compounding evaluations, harness-evolving streams, process-level checkpoints, cross-mechanism continual-learning comparisons, counterfactual shortcut tests, and a first attack on the benchmarks self-improvers trust. The verdicts rhyme: memory helps while it is being written and fades after; plain long context competes with or beats dedicated memory; expanding a harness can make an unchanged model forget; harness evolution beats matched test-time scaling only sometimes; gains shrink about threefold out of distribution; and acceptors, greedy or not, keep worse edits than the proposer offered.
Methods papers have started to absorb the critique. RRSI reports OOD numbers by design 15. ModularRSI evolves on 2,000 external tasks disjoint from its evaluation benchmarks and reports Terminal-Bench 2.0 accuracy rising from 47.57% to 52.43% 34. StarHarness (August 2026) separates proposer-visible search tasks, proposer-hidden selection tasks, and held-out test tasks 33. The memory side has absorbed less: no memory system has yet beaten naive in-context learning on CL-Bench, and agents grading their own memories inflate confident mistakes, which makes a memory system's self-reported improvement exactly the number a protocol should distrust 36. In production, what exists is product-level: A/B deltas, a before-and-after resolution rate, one flywheel, and a survey finding that most "self-improving" deployments disclose too little to check. No production system has published a held-out, matched-budget evaluation of its improvement loop as distinct from its product.
#Open problems
An out-of-distribution standard. RRSI's five held-out benchmarks sit in the same broad domain as its evolution split, and Bad Genius shows that ordinary task holdout can leave benchmark-wide shortcuts intact. No one has agreed whether "out of distribution" means a different benchmark, a different task type, a different model, a different protocol, or a different month of traffic. Without a standard, a threefold shrinkage in one paper and "limited generalization" in another cannot be compared, and every paper can pick the held-out set that flatters it.
Telling compounding from plateau. The RELAI study, whose authors found no earlier protocol for harness optimizers that measured both transfer and continued improvement, tests compounding over just two phases and 22 tasks 1. At lab scale, METR says its data on AI R&D ability cannot distinguish consistent, accelerating, or decelerating improvement 26. The question matters for cost (when to stop spending on search) and for safety, because the claim that a loop is not self-improving past some threshold is a measurement claim that nobody currently knows how to demonstrate.
Cheap estimates of improvement potential. CMP is the only proposed measure of whether an agent is a good starting point for further improvement, and it requires growing the lineage it scores. EvoPathBench's finding that good candidates exist but go unselected suggests the cheaper lever may be selection, not generation, while HarnessOpt-Bench's observation that failure-trace inspection was not associated with held-out gain suggests the obvious proxies may be the wrong ones 21.
The reader should leave with a calibrated prior. When a paper or product says an agent improved itself, the default in September 2026 is that it improved on the tasks that drove the improvement, by an amount that shrinks out of distribution and may not survive the next batch of tasks, that may not beat the same compute spent on retries, and that a greedy acceptor partly invented. Claims that clear a matched-budget baseline, a test set the loop never touched, a second round of new tasks, and a controlled acceptor are rare and worth more. The open question that matters most is whether any agent, memory or harness or weights, can show on a continuous stream that its gains persist and transfer; on SWE-Milestone, AgentCL, CL-Bench, and EvoHarnessBench, none has yet.
#References
● marks sources dated June 24 to September 24, 2026.
- [1]W. Wang, P. Kattakinda, S. Feizi (RELAI.ai), “Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0,” arXiv:2607.14004, 2026-07-15. https://arxiv.org/abs/2607.14004
- [2]Asawa et al., “Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments” (CL-Bench), arXiv:2606.05661, June 2026. https://arxiv.org/abs/2606.05661
- [3]Z. Cheng, J. Xu, H. Chai, J. Sun, P. Viswanath, M. Pan, “AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning,” arXiv:2609.05435, 2026-06-30 (v2 2026-09-21). https://arxiv.org/abs/2609.05435
- [4]D. Lopez-Paz and M. Ranzato, “Gradient Episodic Memory for Continual Learning,” NeurIPS 2017. https://papers.neurips.cc/paper/2017/file/f87522788a2be2d171666752f97ddebb-Paper.pdf
- [5]Z. Ke, V. Patil, H. Shi et al. (Salesforce Research, UNC Chapel Hill, UW–Madison), “EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?,” arXiv:2609.04280, 2026-09-03 (v2 2026-09-10). https://arxiv.org/abs/2609.04280
- [6]Y. Shu, B. Jiménez Gutiérrez, S. P. Jonnalagedda, Y. Yao, H. Sun, Y. Su, “AgentCL,” arXiv:2606.02461, 2026-06-01. https://arxiv.org/abs/2606.02461
- [7]H. Lin, C. Liu, X. Bai, X. Jin, Y. Li, N. Zheng, X. Cao, “Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents” (EvoPathBench), arXiv:2609.24663, 2026-09-21. https://arxiv.org/abs/2609.24663
- [8]W. Wang, P. Piękos, L. Nanbo, F. Laakom, Y. Chen, M. Ostaszewski, M. Zhuge, J. Schmidhuber, “Huxley-Gödel Machine,” ICLR 2026 (oral), arXiv:2510.21614, October 2025. https://arxiv.org/abs/2510.21614
- [9]C.-K. Wu, Z. R. Tam, C.-Y. Lin, Y.-N. Chen, H.-y. Lee, “StreamBench: Towards Benchmarking Continuous Improvement of Language Agents,” NeurIPS 2024, arXiv:2406.08747, June 2024. https://arxiv.org/abs/2406.08747
- [10]T. Wei et al. (UIUC, Google), “Evo-Memory,” arXiv:2511.20857, November 2025 (revised May 2026); see also J. Zheng et al., “LifelongAgentBench,” arXiv:2505.11942, May 2025. https://arxiv.org/abs/2511.20857 ; https://arxiv.org/abs/2505.11942
- [11]D. Yan, J. Liang, D. Hu, R. He, N. J. Yuan, Q. Zhang, T. Tan, “AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?,” arXiv:2608.00155, 2026-07-31. https://arxiv.org/abs/2608.00155
- [12]A. Harrington, N. Saxena, M. Murphy, A. Borovykh, Z. Yun, S. Kamath, A. E. Kyi, T. Darrell, J. Malik, Y. Bai, “When Does Continual Learning Require Learning,” arXiv:2607.07847, 2026-07-08. https://arxiv.org/abs/2607.07847
- [13]Y. Wang et al., “EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective,” arXiv:2605.18421, 2026-05-18 (v2 2026-06-15). https://arxiv.org/abs/2605.18421
- [14]Deng, Chen, Yu et al., “SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution” (first released as EvoClaw), arXiv:2603.13428, March 2026. https://arxiv.org/abs/2603.13428
- [15]P. Xia, R. Han, Z. Wang et al. (Google Cloud AI Research), “Regularized Recursive Self-Improvement” (RRSI), arXiv:2609.24972, 2026-09-21. https://arxiv.org/abs/2609.24972
- [16]S. Lewis, “Same Model, Different Harness: Different Coding-Agent Results,” arXiv:2608.26218, 2026-08-26. https://arxiv.org/abs/2608.26218
- [17]Y. Wang, H. Zhu, Z. Hu et al., “Rethinking the Evaluation of Harness Evolution for Agents,” arXiv:2607.12227, 2026-07-14 (v2 2026-08-27). https://arxiv.org/abs/2607.12227
- [18]Z. Shawn, “PACE: Paired Anytime-valid Commit Evaluation” (anytime-valid acceptance tests for self-evolving agents), arXiv:2606.08106, 2026-06-06. https://arxiv.org/abs/2606.08106
- [19]M. Lin, J. Wu, Z. Wang et al., “Harness Updating Is Not Harness Benefit,” arXiv:2605.30621, 2026-05-28. https://arxiv.org/abs/2605.30621
- [20]G. Zhu, X. Huang, P. Yin, J. Xie, S. Zhang, D. Zhou, “Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts” (CHASE), arXiv:2609.18366, 2026-09-16 (v2 2026-09-18). https://arxiv.org/abs/2609.18366
- [21]V. Ursekar, A. Shanker, Y. Maurya et al. (Scale AI), “HarnessOpt-Bench,” arXiv:2608.06301, 2026-08-06. https://arxiv.org/abs/2608.06301
- [22]Y. Wu, J. Zhang et al. (ByteDance Seed), “HarnessDev,” arXiv:2609.01437, September 2026. https://arxiv.org/abs/2609.01437
- [23]F. Roesner, T. Kohno, “Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks,” arXiv:2609.17817, 2026-09-15. https://arxiv.org/abs/2609.17817
- [24]METR, “Frontier Risk Report,” 2026-05-19, and time-horizon page. https://metr.org/blog/2026-05-19-frontier-risk-report/ ; https://metr.org/time-horizons/
- [25]OpenAI, “Hugging Face model evaluation security incident,” 2026-07-21 (findings 2026-08-26). https://openai.com/index/hugging-face-model-evaluation-security-incident/
- [26]METR, “Summary of METR's predeployment evaluation of Claude Opus 5.5,” 2026-09-22. https://metr.org/blog/2026-09-22-claude-opus-5-5/
- [27]K. Selig, “Most 'self-improving' AI agents don't actually improve,” Serious Tech (Substack), 2026-07-22. https://www.readserioustech.com/p/self-improving-agent-loops-verifier
- [28]A. Luthra, K. Jain, S. Arya, B. Filar, A. Bertiger, “Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis,” arXiv:2608.13608, 2026-08-11. https://arxiv.org/abs/2608.13608
- [29]M. Zhao, “Bugbot now self-improves with learned rules,” Cursor blog, 2026-04-08. https://cursor.com/blog/bugbot-learning
- [30]S. Heule, J. Katz, “Continually improving our agent harness,” Cursor blog, April 2026. https://cursor.com/blog/continually-improving-agent-harness
- [31]Cursor, “Real-time RL for Composer,” Cursor blog, 2026-03-26. https://cursor.com/blog/real-time-rl-for-composer
- [32]A. Shukla et al. (NVIDIA), “Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement,” EACL 2026 Industry Track, arXiv:2510.27051, October 2025. https://arxiv.org/abs/2510.27051
- [33]E. Esakkiraja, D. Akhiyarov, V. Yadav, S. Rajeswar, P. Bechard, S. Nemala, S. Davasam, “StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments,” arXiv:2608.24804, 2026-08-25. https://arxiv.org/abs/2608.24804
- [34]ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement, arXiv:2609.14857, 2026-09-14. https://arxiv.org/abs/2609.14857
- [35]S. Yao, N. Shinn, P. Razavi, K. Narasimhan, “τ-bench,” ICLR 2025, arXiv:2406.12045, June 2024. https://arxiv.org/abs/2406.12045
- [36]“Memory Reward Inflation” (Echo Gap), arXiv:2608.00017, submitted 2026-06-29. https://arxiv.org/abs/2608.00017