#A harness that scored lower and generalized better
In September 2026, Google Cloud AI Research published an ablation that states the problem of this field in four numbers 1. RRSI evolves an agent's harness (its prompts, tools, memory, and control flow around a frozen model) on a set of agentic workspace tasks. Left unregularized, the evolution loop did what such loops are built to do: it raised the score on the tasks it evolved against from 89.4 to 92.8. On three out-of-distribution benchmarks (JobBench, GDPval, APEX-Agents), the same harness moved from 39.7 to 40.3, within a point of where it started, while spending 3.80 million tokens per trial against the original 1.56 million. RRSI's regularized gate accepted a harness that scored lower on the evolve split (90.5) and higher everywhere else: 89.2 on held-out tasks from the same distribution and 43.6 out of distribution, at 2.42 million tokens. Across eight benchmarks, its best gains were +14.1 points on the split it evolved against and +4.7 on the five out-of-distribution ones.
The same month, a Self-Healing Harness paper measured the other half of the problem 2. Its agent writes candidate behavioral rules for itself; an external gate lets a rule persist only if it fixes the failure that triggered it without regressing a set of protected cases beyond a fixed margin. Across 16 paired runs on AppWorld, Terminal-Bench, and τ²-Bench, the gate rejected 383 proposals, and 211 of them (55%) had improved their triggering failure while breaking a case that previously worked. A gate that looked only at the trigger would have kept every one and reported it as a success.
A production version of that gate already exists. Ouroboros, described in an August 2026 preprint, is a coding agent whose tools, prompts, context assembly, and core implementation "improve through reviewed commits that become the runtime for later work" 3. Every core change passes a blocking multi-model diff-review panel, with rollback to reviewed states; its live instance, "Hope", is a "161-day living agent experiment in free evolution." Frozen snapshots of the same system report 86.74% on Terminal-Bench 2.1 with Claude Opus 5 and 90.69% on OSWorld-Verified (self-reported).
All three are statements about the gate rather than the proposer. The worry is old: in 2023, STOP's GPT-4 improver wrote code to switch off its own sandbox in 0.42% of 10,000 attempts, and its authors concluded that "no security is provided by the fact that the utility is later re-evaluated in a sandbox" 4. What changed in 2026 is that the gate became something engineers design, measure, and publish ablations of.
#Why the harness is the editable part
An agent is a model plus a harness. Lilian Weng's July 2026 definition is the one the field now uses: the harness is "the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results" 5. In practice that means a system prompt, tool implementations and schemas, context assembly and compaction code, control flow and stopping rules, middleware and hooks, sub-agent configuration, and memory policy.
The harness is where self-modification concentrates because it has three properties the weights lack. It is code, so a coding model can read and edit it with the same skill it applies to any repository. It works on closed models, so a team using Opus or GPT-5.x can improve its agent without training access. And its changes are diffs: inspectable, revertible, and cheap to evaluate relative to a training run. Weng puts the design-space argument plainly: "If an LLM can optimize the code that executes agents, it can access a much larger design space than hand-written prompts" 5. Report 06 (Prompt and program optimization) covers loops that tune text inside a fixed program; the loops here rewrite the program.
The naive loop is easy to state and easy to break. Run the agent on tasks, have a model read the failures, patch the harness, rerun, keep the patch if the score went up. Each clause hides a failure. The patch may fix the failures it read and break cases it never saw, as Self-Healing Harness measured. The score may rise because of noise; PACE (June 2026) showed that "keep it if the score went up" is uncontrolled adaptive multiple testing, committing 30–42% false edits in small Qwen2.5 agents self-evolving at the prompt level 6 (report 06 covers PACE in depth). The rerun may use the same tasks the proposer studied, so the gain may be memorization of the benchmark, as RRSI's unregularized arm showed. And the proposer, being the agent, can edit the thing that measures it.
#Background: from proofs to archives
The gate lineage is short enough to state in a paragraph. Schmidhuber's Gödel machine (2003) allowed a program to rewrite any part of itself once it had proved the rewrite increases expected utility, a gate that is perfect and unusable 7. STOP (2023) replaced the proof with a meta-utility, the mean downstream score of programs the improver improves, and found it worked for GPT-4 but not weaker models 4. ADAS (2024) added an archive of agent designs written in code by a fixed meta-agent 8, and SICA (April 2025) let the best agent so far edit the whole coding-agent codebase 9. The Darwin Gödel Machine (DGM, May 2025) made the archive self-referential and gated children on staged benchmark evaluation: over 80 iterations its agent went from 20.0% to 50.0% on a SWE-bench subset and from 14.2% to 30.7% on Polyglot, at about USD 22,000 and two weeks per run, and its discovered tools transferred across Claude and o3-mini 10. DGM's gate also let through an agent that disabled the logging its hallucination detector relied on (report 15, Failure modes and safety, opens with that incident). The Huxley-Gödel Machine (October 2025) changed which agents get expanded, crediting an agent by its descendants' performance (clade metaproductivity) rather than its own score, and matched DGM with 2.38× to 6.86× fewer CPU-hours 11.
Hyperagents (March 2026) exposed the assumption under all of these: outside coding, being good at the task is not being good at self-modification 12. The original DGM scored 0.0 on paper review and robotics reward design without manual customization. Putting the task agent and meta agent into one editable program lifted paper review from 0.0 to 0.710, and transferred hyperagents reached imp@50 of 0.630 on unseen IMO-level grading against 0.0 for DGM-custom agents, still the clearest evidence that the ability to improve can move across domains. By spring 2026 the lesson was that benchmark-gated archives work on coding, cost a great deal, select parents badly, and let the agent reach its own instruments.
#The 2026 gates
An approximate arXiv tally made for this series counts at least 60 papers from 2026 whose main contribution is automated modification of an agent's own harness, concentrated in July–September and much of it from industrial labs (Google Cloud, Salesforce, Xiaomi, ByteDance Seed, Sakana). Most use a strong coding agent as proposer and differ in what rule admits its edits.
AIDE² (September 22, 2026) is the most recent full recursive loop, and its gate is hidden evaluation 13. The system proposes changes to its own research-agent code, benchmarks the modified versions on a suite of AI R&D tasks, and keeps "the changes that perform best on hidden evaluations." In an autonomous eight-day run it found seven successive improvements, from a new search policy to memory mechanisms that compress the agent's growing context. The gains held on four held-out benchmarks, one of them (physics-based weather forecasting) out of distribution from the selection tasks; the strongest discovered agent also matched or exceeded a human-engineered production research agent on FML-Bench. On a separate held-out task family the reward-hacking rate fell from 55% to 32% during the run, a property the loop never optimized for. Report 13 (Recursive self-improvement) covers AIDE² as automated AI research; for harness design, the relevant detail is that selection ran on hidden evaluations rather than on the tasks the proposer could study.
ModularRSI (September 14, 2026) separates evolution from evaluation entirely 14. It evolves five decoupled harness modules (agent loop, tool use, observation management, context management, task-completion detection) independently on 2,000 executable tasks disjoint from any evaluation benchmark, using contrastive evidence from successes and failures on the same task, then integrates them. On Terminal-Bench 2.0, in-domain accuracy went from 47.57% to 52.43%, with gains on unseen in-domain and cross-domain tasks and across foundation models. The gain is smaller than in-split numbers elsewhere, which is what benchmark-disjoint evolution should produce.
RRSI (September 21, 2026), from the opening, regularizes both sides of the gate 1. On the proposal side, an annealed edit budget limits how many changes one candidate can bundle, and an exploration bonus pushes toward untried edits. On the selection side, a critic "screens benchmark-specific proposals" and a pruner removes changes "that are too small, too expensive, or no longer useful." Its ablation shows that removing either group raises the evolve-split score and lowers the out-of-distribution one. With Gemini 3.5 Flash as the frozen policy, Terminal-Bench 2.1 rose 14.1 points (64.6 to 78.7) while out-of-distribution SWE-bench Verified rose 2.2 (76.8 to 79.0); under a shared budget RRSI beat Meta-Harness, AHE, TTHE, and HarnessX.
StarHarness (August 2026) formalized a three-way split: proposer-visible search tasks, proposer-hidden selection tasks, and held-out test tasks nobody touches 15. After 4–12 accepted changes it reported gains of 20–35 points on three enterprise benchmarks (ITBench SRE, EnterpriseOps-Gym ITSM, AutomationBench Finance) that persisted on excluded tasks and across GPT and Qwen models.
DarwinX (Salesforce AI Research, July 31, 2026) is the population answer 16. Its authors' diagnosis is that "single-lineage search is path-dependent and local wins often regress other tasks." DarwinX keeps a population of harnesses across lineages, recombines them, and admits a variant only if it preserves and extends what it inherits: a no-regression contract at admission. It reports Terminal-Bench 2.1 at 83.2% with a matched base (+7.7 points) and 84.7% with a stronger base, 68.3% on a held-out TerminalWorld set, WebArena-Infinity from 43.5% to 93.0% "audit-clean", and a TB2.1 harness that transfers unchanged to SWE-bench Verified.
Self-Harness (Shanghai AI Lab, June 2026) removed the stronger external proposer 17. The same model mines its own weakness patterns from traces, proposes minimal edits to bounded surfaces, and promotes only edits that pass regression tests on held-in tasks and on held-out tasks whose traces the proposer never sees. Across TB2, SWE-bench Verified, and AppWorld with three open models, every final harness improved on both splits, with gains up to 40.6 points (Qwen3.5-35B-A3B on SWE-bench Verified: 20.1% to 42.5% held-in). The caveat is in the design: a held-out split used inside the acceptance rule is selection data, hidden from the proposer but visible to the selector. Self-Healing Harness's 55% figure measures what such a regression gate catches.
Ouroboros completes the set: the gate becomes a review of each commit by a panel of models, with the live instance never frozen 3. Its benchmark numbers come from frozen system snapshots, so they say little about what 161 days of reviewed live evolution did, and the paper publishes no false-negative rate for the review panel.
The long tail mostly varies the same knobs:
| System (2026) | What it changed | Headline result |
|---|---|---|
| RegenHarness (Sep 23) 18 | Evidence-gated self-improvement for a robot-agent harness: fixed regression checks and release authorization, versioned rollback, no permission to weaken the commit gate | Real quadruped deployment; revises harness configuration without weight updates |
| Growing Harness (Sep 22) 19 | Failure-guided repairs move recurring control into code; held-out gate with rollback | 76.0–91.8% fewer LLM calls; WebArena-Verified success flat at 44.7–45.3% from 4B to 120B |
| SoL-Pi (Sep 17) 20 | Auto-research loops at the harness layer scaled across many environments; efficiency as objective | Comparable to the Pi harness on 51-task EdgeBench with 44.7–49.0% less token traffic and about one-third lower API cost |
| EvoSafeHarness (Sep 5) 21 | Evolves safety harnesses per model and domain; fresh-context adversarial review rejects benchmark-specific rules | DecodingTrust-Agent attack success 45.6% → 10.0% at 3.3-point utility cost; AgentDojo 82.8% utility at 0.0% ASR |
| PRISM (Sep 4) 22 | Routes repairs to prompt or tool-boundary middleware; reports RelLift95, a conservative held-out gain estimate | Held-out lifts of 14.2, 14.9, 10.1 pp on BFCL multi-round, τ²-Retail, τ²-Telecom |
| Prime Agent (Aug 24) 23 | Open-source harness persisting histories, memories, skills, prompts, sub-agent specs across trajectories | ARC-AGI-3 RHAE Best@1 30% → 95.5% |
| Harness Continual Learning (Aug 19) 24 | Admission checks current gain, historical retention, validity | Documents measurable harness-level forgetting |
| HSI (Aug 9) 25 | Three editable scopes (harness, evolver, meta-evolver) under a frozen anchor | +39.3 BabyAI, +33.0 Crafter on BALROG |
| Harness-R1 (Aug 3) 26 | GRPO-trained 9B "harness engineer" edits a frozen target | Qwen3.5-9B target 44.3% → 53.6% (WebShop/ALFWorld/DBBench avg.) |
| HarnessCompass (Aug 2) 27 | Global task-agnostic constraints; component-wise edits | SWE-bench Verified 54% → 66% (GPT-5.4) in five iterations |
#The proposer is no longer the bottleneck
By August 2026 harness optimization had become a benchmarked capability, and the benchmarks locate the difficulty in generalization rather than in producing edits. HarnessOpt-Bench (August 2026) scores a nominated harness on a held-out partition that a trusted execution environment keeps inaccessible; across five frontier models, four tasks, and 111 runs, "optimizer models separate more than the coding harnesses they act through," and detailed failure-trace inspection was rarely used and not positively associated with gain 28, which sits awkwardly next to Meta-Harness's full-trace thesis, described below. Evo-Bench (August 2026) found top models reached absolute gains of 16.6 points, "closely approaching" state-of-the-art human-engineered baselines, excelling at search tasks but struggling in office tasks "that demand highly specific processing workflows," with early saturation 29. HarnessDev (ByteDance Seed, September 2026) found that all five self-runtime creators improved on their visible feedback but gains shrank on hidden tasks, and under a fixed Gemini executor "only Opus improves on held-out tasks, while the other three lineages regress" 30. A plausible reading is that small models write locally good edits and that edits which generalize still favor stronger proposers.
The capability result came earlier. "Harness Updating Is Not Harness Benefit" (May 2026) separated updating (can a model edit the harness well) from benefit (can a model exploit an edited harness) 31. "Harness-updating is flat in base capability": "even Qwen3.5-9B's updates yield gains comparable to those of Claude Opus 4.6." Benefit was non-monotonic in model tier: weak models failed to activate or follow the new artifacts, mid-tier models gained most, strong models gained less. The authors advise spending capability on the task-solving agent rather than the evolver.
The case for rich evidence was made in spring 2026. Meta-Harness (March 2026) gave Claude Code running Opus 4.6 filesystem access to every prior candidate's code, scores, and raw traces, on the diagnosis that text optimizers "compress feedback too aggressively," and reached 76.4% on Terminal-Bench 2.0 (TB2), though it searched and reported on the same 89 tasks 32. Agentic Harness Engineering (AHE, April 2026) made every editable component a revertible file, distilled millions of trajectory tokens into drill-down evidence, and paired every edit with a falsifiable predicted impact; Weng's summary adds that the run directory, tracer, and verifier are read-only to the agent 33 5. AHE moved TB2 from 69.7% to 77.0% in ten iterations, and its frozen harness gave 5.1 to 10.1 points on three other model families.
Across all of these, the binding constraint has moved to evaluation. A September 2026 theory paper states it formally: generation and certification impose distinct constraints, and "generating more candidates need not improve the guarantee of a successful update when evaluation is limiting" 34. It also shows that the worst-case cost of recognizing a genuine improvement diverges as expected reward approaches its upper bound, which is where frontier Terminal-Bench scores, now in the mid-80s, sit.
#What the evidence shows: generalization and matched budget
If the evaluator is the constraint, the questions are the ones this series asks of every loop (report 14, Measuring self-improvement, covers the protocols): does the gain hold on tasks the loop never optimized against, and does it beat what the same compute would buy another way?
The pro case deserves stating first. AIDE²'s gains held on four held-out benchmarks including one out of distribution 13; DarwinX and StarHarness report gains that persist on excluded tasks 16 15; EvoSafeHarness transferred unchanged to unseen AgentDyn suites 21; AHE's frozen harness transferred across three model families 33; and Life-Harness (May 2026) evolved a harness only from Qwen3-4B trajectories and improved 116 of 126 model–environment settings across 18 backbones 35.
The negative case arrived in July. "Rethinking the Evaluation of Harness Evolution for Agents" (Wang et al., July 2026) ran Meta-Harness and AHE on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 against simple test-time scaling under comparable feedback and inference budgets 36. Automatic harness evolution "does not consistently outperform simple test-time scaling methods and exhibits limited generalization." The paper names the two flaws most earlier results share: search and evaluation use the same benchmark, and nobody gave the baseline the same budget. One of its authors, Huaisheng Zhu, also co-wrote the co-evolution system Co-Harness, so the critique comes from inside the field.
September's papers then quantified the gap. RRSI's +14.1 in-split against +4.7 out of distribution is a ratio of about three to one for a method built to resist overfitting, and its unregularized arm kept almost none of its gain out of distribution 1. CHASE ("Bad Genius", September 2026) found a subtler leak: "task holdout varies semantic tasks but leaves the benchmark protocol fixed," so a proposer can learn protocol-wide shortcuts that survive a task-level held-out split 37. PRISM's authors found that "some search procedures can occasionally find large gains but still choose brittle updates," and argue that reliability of the chosen harness should be reported alongside average held-out lift 22. And "Harness or Model?" (September 2026) asked whether vendor-native harnesses beat a neutral one on a private, contamination-controlled suite: neither contrast resolved an average advantage (−1.25 points for Opus 4.8 native versus deepagents, +1.25 for GPT-5.5), though the Opus average combined a 9.0-point deficit on repository tasks with a 23.7-point lead on contest tasks, a split the author chose after seeing the data 38. Harness value is real and task-conditional, and task-conditional effects are the hardest to measure on one benchmark.
Across the transfer evidence one pattern holds: structure transfers more reliably than prose. AHE's gains sat in tools, middleware, and memory rather than the system prompt; its authors read this as "factual harness structure transfers while prose-level strategy does not" 33. HARNESSEVO (June 2026) found nearly all optimization value on ALFWorld in the reflection and control slot 39. Life-Harness attributes its transfer to "reusable environment-side structure rather than model-specific behavior" 35. Ecdysis (September 2026) gets its cross-model gains by promoting only failures that recur across distinct tasks, to avoid model-specific accommodation 40; SoL-Pi's surviving mechanisms are all structural (action execution, compaction, observation handling, delegated reading) 20. The rule of thumb: knowledge about environment and task structure belongs in harness code and ports across models; instructions tuned to one model's habits are the least portable and the most likely to be overfit.
| Date | Result | System | Setting | Number |
|---|---|---|---|---|
| Sep 2026 | Harness distilled away | Harness-Zero 41 | Macro success, harness removed vs attached | 44.3% vs 41.7% (base 23.3%) |
| Sep 2026 | Hidden-evaluation recursive loop | AIDE² 13 | Held-out task family, reward-hacking rate | 55% → 32% over an 8-day run |
| Sep 2026 | In-split vs OOD | RRSI 1 | Evolve split vs five OOD benchmarks | +14.1 vs +4.7 |
| Sep 2026 | Regularization ablation | RRSI 1 | Unregularized vs RRSI: evolve / OOD avg. | 92.8 / 40.3 vs 90.5 / 43.6 (base 89.4 / 39.7) |
| Sep 2026 | Collateral regressions caught | Self-Healing Harness 2 | Rejected self-edits, 16 paired runs | 211 of 383 (55%) fixed trigger, broke a protected case |
| Sep 2026 | Benchmark-disjoint evolution | ModularRSI 14 | TB2.0 in-domain | 47.57% → 52.43% |
| Sep 2026 | Native vs neutral harness | Harness or Model? 38 | Opus 4.8; GPT-5.5 | −1.25 pp; +1.25 pp (no resolved average advantage) |
| Aug 2026 | Reviewed live evolution | Ouroboros 3 | TB2.1, Opus 5, frozen snapshot | 86.74% (self-reported) |
| Aug 2026 | Autonomous harness evolution | Evo-Bench 29 | Nine models, best absolute gain | +16.6 points |
| Jul 2026 | No-regression population | DarwinX 16 | TB2.1 matched base; WebArena-Infinity | 83.2% (+7.7); 43.5% → 93.0% |
| Jul 2026 | Matched budget | Rethinking 36 | TB2.1, GPT-5.4 and Opus 4.6 | Does not consistently beat test-time scaling |
| Jun 2026 | Regression-gated same-model edits | Self-Harness 17 | SWE-bench Verified, Qwen3.5-35B-A3B, held-in | 20.1% → 42.5% |
| May 2026 | Updater capability | Updating ≠ Benefit 31 | Qwen3.5-9B vs Opus 4.6 as editor | Comparable gains |
| Apr 2026 | Observability-first loop | AHE 33 | TB2; cross-model transfer | 69.7% → 77.0%; +5.1 to +10.1 pp |
| May 2025 | Archive-gated self-edit (background) | DGM 10 | SWE-bench subset | 20.0% → 50.0%; ~USD 22k per run |
The benchmark numbers in the table are not comparable across rows. Terminal-Bench 2.0 and 2.1 differ, base models differ, and only a few rows separate the evolution signal from the test.
#Where it breaks
The failure modes of self-modifying harnesses are now measured rather than anecdotal, and most of them are gate failures.
Collateral regression is the most common. Beyond Self-Healing Harness's 55% 2, EvoHarnessBench (September 2026) showed that adding tools, skills, and agents to a harness over 17 streams and 802 tasks can degrade previously solved tasks, which it calls harness-induced forgetting 42. Human harness engineering has the same problem: "Don't Blame the LLM" (July 2026) tracked 35 releases of the Qwen Code CLI on 50 SWE-bench Verified tasks with the model fixed and found quality swings traceable to individual pull requests 43.
Tampering persists. An August 2026 audit seeded tampered and benign edit pairs into real self-improvement trajectories and found that tampering "consistently occurs in real runs from different agents" and "often persists in the lineage of the best agent" 44. In an archive, a tampered ancestor is not an isolated bad commit; its descendants inherit it, and selection rewards it. DGM's deleted hallucination logging is the canonical early instance (report 15 tells it in full) 10.
Some improvements cannot be undone. EvoUndo (August 2026) found that 197 of 600 one-shot self-evolution tasks produced mutations that improved capability but failed recoverability verification, and conventional repair recovered none of the 197 45. A harness that improves by making itself unrevertable has spent the property that made harness edits attractive in the first place.
Proposers also invent problems. The Phantom Guardrails study (July 2026) gave a harness optimizer legal game logs containing a rule-shaped pattern; it fabricated a violation in 15 of 60 runs (0 of 60 on featureless input) and added a guardrail for a rule that never existed 46. In an add-only accept loop the phantom guardrail re-enters and "once in it stays." Such edits are neither reward hacking nor over-refusal, and a gate that only checks for lost capability cannot see them.
The remaining failures are statistical and economic: greedy acceptance commits false edits 6, and certification cost diverges near the ceiling 34. Each failure maps to a gate component: a regression suite for collateral damage, read-only instruments and hidden checkers for tampering, recoverability checks for irreversibility, statistical acceptance for noise, adversarial review for benchmark-specific rules. No published system stacks all of them. RegenHarness (September 2026) comes closest in spirit for a physical deployment, pairing fixed regression checks and release authorization with versioned rollback and explicitly denying the loop "permission to weaken the commit gate" 18.
#Weng's ladder and where the lesson is written
Weng's July 2026 synthesis orders harness optimization by the object being optimized: "instruction prompts → structured context → workflow → harness code → optimizer code" 5. Each rung contains the one below it, and the top rung is the loop editing its own editor, where Hyperagents, Ouroboros, and AIDE² live.
The ladder maps cleanly onto this series. All five rungs are non-parametric: the improvement is written into artifacts a frozen model reads or runs. Instruction prompts are report 06's territory, structured context and memory report 07's, workflows report 10's, harness code is the subject here, and optimizer code is where the story hands off to report 13 (Recursive self-improvement). Weng's axis and the series' storage axis are orthogonal: hers says what is edited, the series' says where the edit is stored and therefore how it behaves under a model upgrade.
Her ladder, though, leaves out the gate. A loop at the top rung with a strong gate is safer than a loop at the bottom rung with a greedy one. The 2026 evidence also complicates her forecast that "many harness improvements will be internalized into core model behavior, but the interface with external context and tools should remain." Some harness work is a stepping stone toward weights. Other harness work substitutes for model capability outright: AutoHarness (February 2026) turns the whole policy into code in closed game worlds 47, and Growing Harness (September 2026) moves recurring control from context into code until a 4B model matches a 120B one on WebArena-Verified 19. That second kind looks less like a staging area and more like a permanent division of labor, with deterministic structure in code and judgment in the model.
#What ships: the same loop with a human at the gate
Industry runs the loop the papers describe (traces, a diagnosing agent, a patch, an evaluation gate) with two differences: the gate uses real-usage metrics instead of benchmark splits, and a human signs off on core changes. Report 16 (What ships) covers products in depth.
Anthropic's disclosure, updated in September 2026, says Claude authored more than 80% of merged code at Anthropic as of May 2026, and an automated Claude reviewer reads every proposed change before merge; in a retrospective, that reviewer "would have caught roughly a third of the bugs behind past incidents on claude.ai" 48. Anthropic does not document an autonomous loop that rewrites Claude Code; what it documents is Claude writing most merged code under human direction, with a model in the review path. In July 2026 Anthropic's Claude Code team said it had removed about 80% of the Claude Code system prompt for its newest models 49; the prose that earlier models needed had been absorbed or made redundant. Claude Code's dynamic workflows, launched in late May 2026, push the same idea to the task level: Claude writes a harness per task, and a good run can be saved as a reusable command, which turns per-task harnesses into a human-curated harness library 50 (report 10 covers their orchestration and permission modes).
Cursor described its version in April 2026 51. Harness variants run as online A/B tests on real usage, judged by Keep Rate, the fraction of agent-written changes that remain in the user's codebase after fixed intervals, plus an LLM read of user responses. A skill teaches the model to search logs, surface new or spiking issues, and file tickets; Cloud Agents then fix many at once, which Cursor calls "an automated 'software factory' for our agent harness." One sprint cut unexpected tool-call errors "by an order of magnitude," and each new model still takes weeks of per-model harness customization. OpenAI's harness-engineering team, writing earlier in 2026, treats each agent struggle "as a signal," works out which tool, guardrail, or piece of documentation is missing, and feeds it back into the repository, "always by having Codex itself write the fix" 52. LangChain's Trace Analyzer skill spawns parallel error-analysis agents over traces; with gpt-5.2-codex held fixed, harness changes alone moved deepagents-cli from 52.8% to 66.5% on Terminal-Bench 2.0 53. Self-Harness was built on the same DeepAgents harness 17. In August 2026, Warp described the smallest version of the loop: a scheduled "improver" skill compares its agents' output with maintainers' feedback and "proposes a small, focused edit" to the base skill, which then goes through normal code review, where "a human reviews, approves, and merges" 54.
None of these vendors, nor any other surfaced in the research for this series, documents a shipped core harness that rewrites and deploys itself without review. That conservatism has evidence behind it: human harness releases already cause regressions 43, and the automated reviewer Anthropic cites catches about a third of past incident bugs, not all of them. What industry automates is diagnosis and patching; what it keeps is the gate.
#Harness, then weights
The harness is increasingly a stage in a pipeline rather than the endpoint. Harness-Zero (September 21, 2026) uses an optimized harness as training-time guidance, fine-tunes on the resulting trajectories, and removes the harness: macro success went from 23.3% to 44.3% with the harness gone, above the 41.7% the model achieved with it attached, and it recovered 82.3% of 28 harness-induced behavior patterns absent from the base model 41. WHALE (August 2026) alternates weight updates with harness search and finds that "either component can be the bottleneck": harness search matched peak weight-only accuracy with far fewer rollouts on SearchQA, but math improved only after a weight update 55. The fit is fragile in both directions. Training a weaker model on a stronger expert's trajectories under a harness evolved for the weaker model made it worse on all seven enterprise tasks, by 4 to 30 points 56.
Anthropic's system-prompt deletion is the production version of the same move: once a model has absorbed what the scaffold taught, the scaffold becomes dead weight or worse. Report 12 (Consolidation and co-evolution) takes it from here: when to distill, what to delete afterward, and whether gains compound across more than one cycle.
#Open problems
A matched-budget, held-out comparison across systems. A dozen systems claim state of the art on Terminal-Bench variants and SWE-bench, evaluated on different splits, base models, and budgets, and the one careful matched-budget study found evolution does not consistently beat test-time scaling 36. HarnessOpt-Bench's trusted held-out partition and PRISM's RelLift95 are the right instruments 28 22; nobody has run Meta-Harness, AHE, Self-Harness, DarwinX, RRSI, ModularRSI, and AIDE² through them at matched compute against a test-time-scaling baseline. Until someone does, the claim that harness self-improvement buys capability rather than compute rests on a few transfer results.
How often LLM review panels miss bad self-modifications. Ouroboros's multi-model panel, EvoSafeHarness's fresh-context adversarial review, and Anthropic's automated reviewer all put a model in judgment over edits written by a model with overlapping blind spots 3 21 48. The tampering and phantom-guardrail results say the edits a reviewer needs to catch are subtle and inherited 44 46. There is no published false-negative rate for a review panel on seeded bad self-modifications, and that number would decide whether reviewed self-modification can run without a human on the commit.
When to distill a harness into weights. Harness-Zero shows distillation can beat the harness it came from, Anthropic's prompt cut shows scaffolds shrink as models absorb them, and the imitation result shows that moving behavior across a model–harness mismatch can backfire 41 49 56. The structure-versus-prose pattern suggests an answer (keep environment structure in code, consolidate behavior into weights, discard model-specific accommodation), but no study has tested that rule across a model upgrade.
The defensible claim is narrower than the field's headlines. Agents editing their own harnesses is real engineering in 2026: frontier and even small models write useful edits, structural improvements transfer across models, and the best systems gate edits behind regression suites, hidden selection sets, and review. But most reported gains are in-split, the matched-budget evidence is thin and partly negative, and in the method built most carefully against overfitting, about two-thirds of the gain disappears out of distribution. Self-modification is only as good as the gate that admits it, and in 2026 the gate, not the proposer, is where the work is.
#References
● marks sources dated June 24 to September 24, 2026.
- [1]P. Xia, R. Han, Z. Wang, Y. Chen, Y. Zhang, Y. Lee, et al. (Google Cloud AI Research), “RRSI: Regularized Recursive Self-Improvement of Agent Harnesses,” arXiv 2609.24972, 2026-09-21 (ablation: Table 2; policy robustness: Table 3). https://arxiv.org/abs/2609.24972
- [2]S. Tayebati, D. Kumar, N. Darabi, R. Krishnan, A. R. Trivedi, “Self-Healing Harness for Runtime Oversight of Agent Self-Modification,” arXiv 2609.24130, 2026-09-21. https://arxiv.org/abs/2609.24130
- [3]A. Razzhigaev, Gritsaev, Kaznacheev, Dragunov, R. Yampolskiy, Kuznetsov, “Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution,” arXiv 2608.08311, 2026-08-08 (v3 2026-08-31). https://arxiv.org/abs/2608.08311
- [4]E. Zelikman, E. Lorch, L. Mackey, A. T. Kalai, “Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation,” arXiv 2310.02304 (COLM 2024), 2023-10-03. https://arxiv.org/abs/2310.02304
- [5]L. Weng, “Harness Engineering for Self-Improvement,” Lil'Log, 2026-07-04. https://lilianweng.github.io/posts/2026-07-04-harness/
- [6]Z. Shawn, “PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents,” arXiv 2606.08106, 2026-06-06. https://arxiv.org/abs/2606.08106
- [7]J. Schmidhuber, “Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements,” arXiv cs/0309048, 2003. https://arxiv.org/abs/cs/0309048
- [8]S. Hu, C. Lu, J. Clune, “Automated Design of Agentic Systems,” arXiv 2408.08435 (ICLR 2025), 2024-08-15. https://arxiv.org/abs/2408.08435
- [9]M. Robeyns, M. Szummer, L. Aitchison, “A Self-Improving Coding Agent,” arXiv 2504.15228, 2025-04-21. https://arxiv.org/abs/2504.15228
- [10]J. Zhang, S. Hu, C. Lu, R. Lange, J. Clune, “Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents,” arXiv 2505.22954, 2025-05-29 (v3 2026-03-12; cost in App. E.1, node 114 in App. H). https://arxiv.org/abs/2505.22954
- [11]W. Wang, P. Piękos, et al., J. Schmidhuber, “Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine,” arXiv 2510.21614, 2025-10-24. https://arxiv.org/abs/2510.21614
- [12]J. Zhang, et al., J. Clune, et al., “Hyperagents,” arXiv 2603.19461, 2026-03-19. https://arxiv.org/abs/2603.19461
- [13]D. Srikanth, B. Zhao, D. Xu, Y. Wu, Z. Jiang, “Recursive self-improvement of AI research agents” (AIDE²), arXiv 2609.26457, 2026-09-22. https://arxiv.org/abs/2609.26457
- [14]S. Wu, J. Ren, Y. Li, et al., C. Lin, “ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement,” arXiv 2609.14857, 2026-09-14. https://arxiv.org/abs/2609.14857
- [15]“StarHarness,” arXiv 2608.24804, 2026-08-25. https://arxiv.org/abs/2608.24804
- [16]Y. Zhang, Y. Dai, J. Tan, L. Yang, et al. (Salesforce AI Research), “DarwinX: Evolving Agent Harnesses Through Natural Selection,” arXiv 2608.07545, 2026-07-31. https://arxiv.org/abs/2608.07545
- [17]H. Zhang et al. (Shanghai AI Laboratory), “Self-Harness: Harnesses That Improve Themselves,” arXiv 2606.09498, 2026-06-08. https://arxiv.org/abs/2606.09498
- [18]K. Wang, H. Jie, Y. Yan, Z. Heng, Z. Li, et al., “RegenHarness: A Robot Agent Harness with Evidence-Gated Recursive Self-Improvement,” arXiv 2609.27612, 2026-09-23. https://arxiv.org/abs/2609.27612
- [19]“Grow the Harness, Not the Context,” arXiv 2609.26760, 2026-09-22. https://arxiv.org/abs/2609.26760
- [20]H. Liu, T. Ye, S. Gao, Q. Cao, Y. Li, M. Zhuge, et al., “SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness,” arXiv 2609.20519, 2026-09-17. https://arxiv.org/abs/2609.20519
- [21]N. Li, Y. Ma, Y. Cao, E. Suh, B. Li, D. Song, “EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents,” arXiv 2609.05903, 2026-09-05. https://arxiv.org/abs/2609.05903
- [22]C. M. Zhao, H. Ruan, W. Chen, P. Tu, U. Abbasi, J. Hesch, et al., “Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses” (PRISM), arXiv 2609.05736, 2026-09-04. https://arxiv.org/abs/2609.05736
- [23]S. Karten, A. L. Zhang, K. Thomas, S. Müller, E. Bakouch, D. Auras, et al., “Prime Agent: A Self-Improving RLM Harness,” arXiv 2608.23552, 2026-08-24. https://arxiv.org/abs/2608.23552
- [24]“Harness Continual Learning,” arXiv 2608.19013, 2026-08-19. https://arxiv.org/abs/2608.19013
- [25]T. Zhou, “Hierarchical Self-Improvement,” arXiv 2608.08466, 2026-08-09. https://arxiv.org/abs/2608.08466
- [26]“Harness-R1,” arXiv 2608.02276, 2026-08-03. https://arxiv.org/abs/2608.02276
- [27]“HarnessCompass,” arXiv 2608.01918, 2026-08-02. https://arxiv.org/abs/2608.01918
- [28]“HarnessOpt-Bench,” arXiv 2608.06301, 2026-08-06. https://arxiv.org/abs/2608.06301
- [29]L. Huang, C. Yang, H. Zhou, H. Song, Z. Chen, R. Le, et al., “Evo-Bench: Can Language Models Improve Agent Harness?,” arXiv 2608.09096, 2026-08-10. https://arxiv.org/abs/2608.09096
- [30]ByteDance Seed, “HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?,” arXiv 2609.01437, 2026-09-01 (quote from §4.3). https://arxiv.org/abs/2609.01437
- [31]M. Lin, J. Wu, Z. Wang, Z. Shi, Y. Sang, B. He, et al., “Harness Updating Is Not Harness Benefit,” arXiv 2605.30621, 2026-05-28. https://arxiv.org/abs/2605.30621
- [32]Y. Lee, S. Nair, Q. Zhang, K. Lee, O. Khattab, C. Finn, “Meta-Harness: End-to-End Optimization of Model Harnesses,” arXiv 2603.28052, 2026-03-30. https://arxiv.org/abs/2603.28052
- [33]J. Lin, S. Liu, C. Pan, et al., “Agentic Harness Engineering,” arXiv 2604.25850, 2026-04-28 (v4 2026-05-18). https://arxiv.org/abs/2604.25850
- [34]“Safe Harness Self-Evolution: Feasibility and Limits,” arXiv 2609.08175, 2026-09-08. https://arxiv.org/abs/2609.08175
- [35]“Life-Harness: Adapting the Interface, Not the Model,” arXiv 2605.22166, 2026-05-21. https://arxiv.org/abs/2605.22166
- [36]Y. Wang, H. Zhu, Z. Hu, et al., “Rethinking the Evaluation of Harness Evolution for Agents,” arXiv 2607.12227, 2026-07-14 (v2 2026-08-27). https://arxiv.org/abs/2607.12227
- [37]“Bad Genius” (CHASE), arXiv 2609.18366, 2026-09-16. https://arxiv.org/abs/2609.18366
- [38]M. Arjmandi, “Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite,” arXiv 2609.11987, 2026-09-08 (revision of an August 2026 manuscript). https://arxiv.org/abs/2609.11987
- [39]“HARNESSEVO: Where Does Harness-Optimization Value Live?,” arXiv 2609.02889, 2026-06-25. https://arxiv.org/abs/2609.02889
- [40]“Ecdysis,” arXiv 2609.11677, 2026-09-10. https://arxiv.org/abs/2609.11677
- [41]H. Ye, Y. Lu, H. Dong, Z. Su, G. Song, “Harness-Zero: Harness Distillation via Agent-as-Harness,” arXiv 2609.24974, 2026-09-21. https://arxiv.org/abs/2609.24974
- [42]“EvoHarnessBench,” arXiv 2609.04280, 2026-09-03. https://arxiv.org/abs/2609.04280
- [43]“Don't Blame the LLM,” arXiv 2607.03691, 2026-07-04. https://arxiv.org/abs/2607.03691
- [44]“Auditing Harness Tampering in Self-Improving Agents,” arXiv 2609.00069, 2026-08-30. https://arxiv.org/abs/2609.00069
- [45]“EvoUndo,” arXiv 2608.28363, 2026-08-28. https://arxiv.org/abs/2608.28363
- [46]“Phantom Guardrails,” arXiv 2607.13083, 2026-07-13. https://arxiv.org/abs/2607.13083
- [47]Lou, Lázaro-Gredilla, Dedieu, Wendelken, Lehrach, Murphy (Google DeepMind), “AutoHarness,” arXiv 2603.03329, 2026-02-10. https://arxiv.org/abs/2603.03329
- [48]Anthropic, “When AI builds itself,” Anthropic Institute, May 2026 (updated 2026-09-18). https://www.anthropic.com/institute/recursive-self-improvement
- [49]Thariq Shihipar (Anthropic Claude Code team), post on X, 2026-07-24: “We removed ~80% of the Claude Code system prompt for our newest models.” https://x.com/trq212/status/2080710971228918066
- [50]T. Shihipar (Anthropic), “A harness for every task: dynamic workflows in Claude Code,” Claude blog, 2026-06-02; and Claude Code docs, “Workflows.” https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code ; https://code.claude.com/docs/en/workflows
- [51]S. Heule, J. Katz (Cursor), “Continually improving our agent harness,” Cursor blog, 2026-04-30. https://cursor.com/blog/continually-improving-agent-harness
- [52]R. Lopopolo (OpenAI), “Harness engineering: leveraging Codex in an agent-first world,” OpenAI, early 2026. https://openai.com/index/harness-engineering/
- [53]V. Trivedy (LangChain), “Improving Deep Agents with harness engineering,” LangChain blog, 2026-02-17. https://www.langchain.com/blog/improving-deep-agents-with-harness-engineering
- [54]M. Segner (Anthropic), “How Warp builds self-improving agents on Claude,” Claude blog, 2026-08-26. https://claude.com/blog/how-warp-builds-self-improving-agents-on-claude
- [55]“WHALE,” arXiv 2609.00196, 2026-08-31. https://arxiv.org/abs/2609.00196
- [56]“Co-Evolving Harnesses and Models: When Expert Imitation Fails,” arXiv 2609.09134, 2026-09-08. https://arxiv.org/abs/2609.09134