Research series · 2026 · Report 12 of 16
Report 12 of 16 · Part III · Bridge · Both

Consolidation and co-evolution

where should learning live?

The long argument over whether agents should learn in their context or in their weights has an emerging answer: both, at two speeds. Systems that optimize harness and weights together find that each lever unlocks the other, and the summer of 2026 showed small interleaved updates beating a single harness-then-weights handoff; when learning moves from context into weights, the distillation objective decides how much survives; code-level harness fixes survive a model swap while prose does not; and labs now co-design model and harness, then delete the scaffolding the model has absorbed. The evidence is a few cycles deep and has hygiene problems, so the design is settled faster than the proof.

ContentsReport 12 · Bridge
$ tree ./12-consolidation-coevolution
./12-consolidation-coevolution
├── 01-a-model-that-beat-its-own-harness# 331 words
├── 02-two-camps-and-a-map-from-neuroscience# 526 words
├── 03-optimizing-harness-and-weights-together# 742 words
├── 04-moving-learning-into-weights-without…# 806 words · 1 figure
├── 05-prompts-first-then-gradients-then-both# 240 words
├── 06-what-the-evidence-shows-and-how-far-it…# 575 words
├── 07-what-survives-a-model-swap# 468 words
├── 08-labs-train-the-pair-then-delete-the…# 400 words
├── 09-weights-that-write-context# 296 words
├── 10-where-learning-should-live# 517 words · 1 figure
├── 11-state-of-play-and-what-ships-2026# 270 words
└── 12-open-problems# 447 words
12 sections · 47 references · 2 figures
23.3% → 44.3%
task success after the model absorbed its harness
64.8% vs 3.8%
in-context gain kept by forward KL vs SFT
5.32 / 9.16
points small alternating updates add over stagewise
47
references · 19 from Jun 24 – Sep 24, 2026
23 min
reading time
Storage
Both: a versioned harness/context store (fast) and low-rank adapters (LoRA) or full weights (slow); trained controllers whose weights write context for frozen models
Engines
LLM critics editing harness code and prompts; reinforcement learning (RL) and supervised fine-tuning (SFT); distillation from a context- or harness-conditioned teacher; RL-trained memory and harness policies
Evaluator
Task graders, unit tests, verifiers; held-in/held-out acceptance gates; few matched-budget baselines
Loop timescale
Harness: per episode or generation. Weights: interleaved (WHALE), on plateau (SIA), every ~5 hours (Cursor), per model generation (labs)
Loop closure
Research loops mostly closed with human-designed gates; lab co-design is human-on-the-loop
Evidence maturity
Early: small models (2–9B) or at most two cycles at larger scale, a handful of tasks, mostly preprints; lab evidence is disclosure, not controlled study

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 12

#A model that beat its own harness

On September 21, 2026, Ye, Lu, Dong, Su and Song posted Harness-Zero, a recipe for training a harness into a model and then throwing the harness away 1. They evolved a specialized harness for each domain (knowledge work, tool use, science) using Kimi K3 under Kimi Code. They then ran a Qwen3.5-9B student inside a deliberately bare target harness, a minimal mini-SWE-agent-style loop with a fixed system prompt and one Bash tool, while GPT-5.6 Sol, guided by the evolved harness, corrected the student's responses before they executed. Fine-tuning on those corrected trajectories produced a model that, with the specialized harness removed, raised macro-average task success from 23.3% to 44.3%. With the specialized harness still attached, the untrained base model reached 41.7%. The distilled model recovered, on average, 82.3% of 28 harness-induced behavior patterns that the base model lacked 1.

The result is a week-old preprint on one student model, and the teacher in the loop is a frontier model, so part of what got distilled is GPT-5.6 Sol's judgment rather than the harness alone. The result still makes the storage question concrete. An evolved harness is expensive to build and tied to one domain; a general-purpose agent either settles for a mediocre shared harness or routes among a growing zoo of specialized ones. Harness-Zero's answer is to use the specialized harness as training-time guidance, write what it induces into the weights, and ship one plain harness. The harness did the discovery; the weights did the keeping.

Every self-improving system has to decide where its improvements are written. Written into text (prompts, memory, skills, harness code), learning is cheap, legible, reversible, and available on closed models. Written into weights, it is compact, fast at inference, and able to change what the model can do rather than what it is told. The hard question is no longer which store wins. The 2026 evidence is about the machinery that uses both, moves lessons from one to the other, and decides which lessons belong where.

Section 02 / 12

#Two camps and a map from neuroscience

The sharpest statement of the context-first position comes from Letta, which builds stateful agents. In June 2026 it wrote: "While models are a depreciating asset, learning and memory represented in token space is an appreciating asset" 2, sharpening its December 2025 line: "The weights are temporary; the learned context is what persists" 3. The reasoning is practical. Most builders use closed models whose weights they cannot touch, a new model generation arrives every few months, and a memory file or skill library carries over to the next model while a fine-tune does not. Letta's roadmap calls for "memory models" trained with RL to write durable context "that transfers across model generations" 2 (see 07, Experiential memory).

The consolidation camp argues that context is a staging area. In May 2026 Dwarkesh Patel, relaying Andrej Karpathy, put it this way: "Our working memory gets wiped regularly. What we actually have is a consolidation process (sleep) that distills stuff into the brain, in a weird and lossy way… It's not obvious it's something you can get for free from doing long enough RL loops" 4. Within eight days two papers put sleep in their titles. One converts recent context into persistent state-space fast weights through offline recurrent passes and then clears the KV cache 5; the other adds a "sleep" stage in which a model consolidates memories by distillation and "dreams" an RL-generated curriculum to rehearse new knowledge 6.

Both camps now concede a hybrid. Letta grants that "the primary advantage of parametric memory is efficiency" and that distilling memories into weights "can provide additional gains in personalization and efficiency" when the weights are accessible 3 2. The sleep papers treat context as the place where new information lands first. The disagreement is over which store is primary and how often the slow one should be written.

Complementary Learning Systems (CLS) theory gives the structure a name. McClelland, McNaughton and O'Reilly argued in 1995 that the brain needs a fast hippocampal store that records episodes after one exposure and a slow neocortical store that learns structure through interleaved replay, because one store cannot do both without catastrophic interference 7. The mapping to agents is direct. Context, memory, skills and harness are the fast store: written every episode, specific, cheap to revise. Weights are the slow store: updated rarely, general, expensive to revise, and prone to interference when updated carelessly. Consolidation is replay from the first into the second. The analogy's one testable prediction is that the slow store should learn from what the fast store computes, not from raw experience; a 2025 DeepMind study co-authored by McClelland found that fine-tuning generalized better when its data included the model's in-context reasoning traces 8. By August 2026 engineers were building CLS into agent memory by name: Dual-Layer Agentic Memory routes incoming facts into an external store and periodically consolidates high-value entries into the weights with supervised fine-tuning; a 1.7B/8B routing cascade pruned up to 68% of redundant external memory while keeping over 98% of the QA exact match of a keep-everything baseline 9. That is enough work for the analogy. The engineering question is how a loop should switch between the two stores.

Section 03 / 12

#Optimizing harness and weights together

The first systems to put both levers inside one loop arrived between May and September 2026, and they share a shape. The harness is edited while weights are held fixed, the weights are trained under the improved harness, and each change exposes a new bottleneck in the other. What changed over the summer is the answer to how often to switch.

WHALE, posted August 31 by Kim, Lee, Finn and Lee, is the cleanest test 10. It alternates online rejection-sampling fine-tuning with Meta-Harness search (Claude Opus 4.7 as the harness proposer), using Qwen3.5-2B and 4B agents on search QA, math and chess puzzles. It beat weight-only, harness-only and Fast-Slow Training baselines by 4.15 to 24.38 points in best mean@8 accuracy. The per-domain detail matters most: "Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update." And small interleaved updates beat a stagewise schedule that completes all weight optimization before beginning harness search, by 5.32 and 9.16 points in the two domains the authors report, at lower rollout cost 10. At this scale the answer to "harness or weights first" is neither: switch often.

Earlier systems switched once or twice. SIA, from May, lets a Claude Sonnet 4.6 Feedback-Agent choose after each generation between rewriting a gpt-oss-120b agent's scaffold and running RL on its LoRA adapter; in practice it ran harness search until progress stalled and then switched 11. On LawBench's 191-class charge prediction, top-1 accuracy went from 13.5% to 50.0% after harness search to 70.1% after weight updates, against a prior best of 45.0%. On a single-cell denoising task, scaffold rewrites carried the score from 0.048 to 0.241 and stalled; the first weight-update checkpoint then wrote a two-line np.clip plus np.rint step that rounds imputed counts to non-negative integers, "a biological invariant that is trivially correct yet absent from any prior scaffold version," lifting the score to 0.289 11.

Co-Harness, from July, adds explicit acceptance rules 12. An LLM "HarnessCritic" labels each failed trajectory with a root cause and proposes a local diff, which is kept only if it improves a held-in slice showing the failure without regressing a held-out slice; accepted diffs go into a versioned registry "that serves as an audit trail and enables full rollback." Verifier-passing trajectories under the accepted harness then fine-tune the next model. On tool-integrated math with Qwen3-8B and Qwen3-32B, average pass@1 across AIME24, AIME25 and HMMT25 moved from 54.2 with a human harness to 58.5 with the evolved harness, then 73.5 and 78.9 after two rounds of fine-tuning. Its 200-hour case study shows what the harness lever buys: the initial configuration crashed after 32 seconds on a vLLM KV-cache misallocation, and over 22 versions the critic fixed that, cured a zombie-thread out-of-memory failure, and found an 8.7× batching speedup. It also added domain prompts that cost about 16.5 points by clashing with Qwen3's trained reasoning format, and the registry rolled them back. No amount of RL on a trainer that crashes in 32 seconds produces a gradient.

Other August systems make the harness an explicit data engine. HELIX, from HKU, ran one evolution round on code repair: a 65-candidate harness portfolio found a fixed harness that improved task coverage by 4.0% over the Pi agent, and a 200-slot slice of sibling runs yielded 438 verified SFT, critic, filter and preference records for the next model 13. Macaron-V1, Mind Lab's August technical report, is the production-scale design: each release is a versioned model–harness pair (a frozen base, four LoRA specialists selected per turn, and a TOML "Harness Context Protocol" runtime configuration), new capabilities appear first as harness artifacts, and "when trajectory evidence warrants a parameter update, the change can then be transferred into a specialist adapter" 14. Its authors are candid that "the present evaluation measures the configuration-search stage, not repeated transfer across adapter generations." EvoTrainer, from June, applies the idea to the RL training harness itself, reporting in its abstract that it matches or exceeds human-engineered RL recipes on math, competitive programming and repository-level SWE 15.

The mechanism is consistent across these papers. The harness fixes failures that are operational rather than cognitive, and those failures bound what any training signal can reach. Weight updates need decent rollouts to learn from, which the harness supplies. And the weights then reach behaviors the harness cannot express, which change which harness is best, which is why WHALE's frequent switching wins.

Section 04 / 12

#Moving learning into weights without losing it

Once a loop decides a lesson should move into weights, the next question is how. Every method is a variant of context distillation, an operator dating to 2022: run the model with privileged context as a teacher, train it without that context as a student 16. The variants differ in who samples the trajectories and what the teacher sees, and in 2026 those differences turned out to be enormous.

Experience Distillation, a July 2026 preprint, gave a model long histories of its own prior attempts at a task (60 to 600 turns, more than 80k tokens) and measured how much of the in-context gain survives once the history is removed and the lesson trained in 17. A teacher conditioned on the experience generated one-step decisions at each recorded point, and the student was trained on them with teacher-sampled forward KL. That kept 64.8% of the in-context gain on 749 SWE tasks and 93.4% on six TaleSuite text games. Plain supervised fine-tuning on the same experience kept 3.8% and −2.6%. On two games, on-policy distillation with reverse KL, which much of the field treats as the better default (see 04, Self-training and distillation), kept 9.1% and 0.4% at matched compute, against 96.7% and 98.4% for forward KL, because "without the experience context, the student rarely samples trajectories containing the high-quality decisions." The caveat is scope: the paper measures retention on the same tasks the experience came from, not transfer to new ones, and the base models are in-house.

Fig 12.1 · How much of the lesson survivesExperience Distillation · Jul 2026
SHARE OF IN-CONTEXT GAIN RETAINED IN WEIGHTS (%) 0 100 SWE · 749 TASKS Forward KL, teacher-sampled 64.8 SFT on the same experience 3.8 TALESUITE · SIX TEXT GAMES Forward KL, teacher-sampled 93.4 SFT on the same experience −2.6 TWO GAMES · MATCHED COMPUTE Forward KL 96.7 98.4 Reverse-KL on-policy distillation 9.1 0.4
The objective decides what survives. Each bar is the share of the in-context gain a model keeps once its long experience history is removed and the lesson is trained into weights. A teacher that sees the experience and trains the student with teacher-sampled forward KL keeps most of the gain; plain SFT on the same experience keeps almost none, and on two games reverse-KL on-policy distillation also loses it at matched compute, because the student rarely samples the good decisions 17. Retention is measured on the same tasks the experience came from, not on new ones 17.

September's results locate the rule more precisely. Yu et al. first evolved a harness for a weaker model on seven enterprise agent tasks, then trained the weaker model on a stronger expert's complete trajectories under that harness 18. Performance regressed on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, though the same imitation helped under the unevolved harness. The weaker model "adopts the expert's planning strategy without the competence to execute it and no longer matches the harness evolved around its native planning style." Their fix keeps the student's own rollout, localizes the failing turn, and asks the expert to rewrite only that turn. Harness-Zero's correction-before-execution works the same way 1, as does Cursor's Composer 2.5, which in May began inserting a short hint at a turn where the model went wrong (for a hallucinated tool call, a "Reminder: Available tools…" note) and distilling the hinted distribution into the unhinted model for that turn alone 19. Taken with Experience Distillation, the lesson is that consolidation targets must be dense and informed by the context, but they must sit on states the student visits itself and match the harness it will run in. Raw imitation of someone else's trajectories fails the second test; student-sampled reverse KL can fail the first.

The bandwidth argument explains why distillation beats outcome RL for this job. Thinking Machines argued in 2025 that policy-gradient RL delivers "only O(1) bits per episode" while supervised targets carry information on every token 20. A context-conditioned teacher turns a rich context into token-level targets that encode what the context changed about each decision; an outcome reward carries a bit or two.

Consolidation also needs a stopping rule, and in September one appeared. RetireOPD trains a skill-free student with RL plus on-policy distillation from a skill-conditioned teacher, and lets the student drop the teacher "once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate," continuing with RL alone 21. Across Qwen2.5 models from 1.5B to 7B it improved ALFWorld success over an RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpassed its own teacher in every setting. The authors also found that "privileged information alone does not always make a teacher reliable," so the teacher is itself trained against environment rewards first. Retiring the teacher is scaffold deletion at training time.

The last failure mode is overfitting to one harness. Alibaba's Taobao Live team found in August that fine-tuning a compact model under a fixed harness configuration lowered IFEval by 7.7 points; training instead under task-preserving transformations of skill identifiers and content, tool schemas, prompt structures and hook functions, followed by on-policy distillation and RL in the augmented environments, avoided the regression (83.5 on IFEval) and scored 94.6 on a harness-variant QA set against the base model's 75.4, in a system deployed with positive online A/B results 22. Consolidating a harness into weights should make a model better inside that harness without making it brittle outside it.

Two older results bracket these. SEAL (2025) let a model write its own fine-tuning data and showed both the promise of self-directed consolidation and its cost in forgetting under sequential edits 23; Cartridges (2025) showed a trained KV cache can match in-context learning on a document at 38.6× less memory 24. Neither changes the 2026 picture.

Section 05 / 12

#Prompts first, then gradients, then both

A hybrid only makes sense if the levers compose. If RL simply relearned what prompt optimization found, or overwrote it, the second step would be wasted work. The direct experiments in the DSPy lineage (see 06, Prompt and program optimization) find composition. mmGRPO, revised in May 2026, reported that averaged over three seeds, two 8B models and three tasks, vanilla chain-of-thought scored 66.3, MIPROv2 prompt optimization 70.0, multi-module GRPO alone 71.2, and prompt optimization followed by GRPO 73.4 25; the authors attribute the gain to "the value of high-quality rollouts at the start of mmGRPO training," and the prompt step cost 1.4 hours on one H100 against 18.7 hours on two for RL. E-SPL, from February, samples each RL batch under several evolving system prompts and moved an AIME-trained model's success on BeyondAIME from 38.8% with RL alone to 45.1%, where prompt evolution alone reached 40.0% 26.

Until August, every composition result ran prompts before weights, and the reverse order (optimizing a harness for a model that has already been trained) was untested. WHALE is the first controlled look: its math harness gains appeared only after a weight update 10. That is on 2B models, and at frontier scale the question remains open, though it is the situation every lab is in when it ships a new model into an existing harness. Co-Harness's clashing domain prompts are the small-scale warning that a context instruction can fight a trained-in behavior 12.

Section 06 / 12

#What the evidence shows, and how far it goes

Date System What moves Result Baseline and context Caveat
Sep 2026 Harness-Zero 1 Evolved harness → Qwen3.5-9B weights, harness removed Macro success 44.3% Base 23.3%; base with harness attached 41.7% GPT-5.6 Sol teacher in the loop; one student; preprint
Sep 2026 RetireOPD 21 Skill context → weights, teacher retired ALFWorld +14.1–18.8%, WebShop +11.8–19.0% RL baseline; beats own teacher Qwen2.5 1.5–7B
Sep 2026 Yu et al. 18 Expert trajectories under evolved harness → weaker model −4 to −30 points on all 7 tasks Same imitation helps under unevolved harness Enterprise tasks; fixed by on-policy turn correction
Aug 2026 WHALE 10 Interleaved harness search and RFT +4.15 to +24.38 pts best mean@8 Weight-only, harness-only, Fast-Slow Qwen3.5-2B/4B, three domains
Aug 2026 HAT (Taobao Live) 22 Harness-augmented training Harness-variant QA 94.6; IFEval 83.5 Base 75.4; fixed-harness SFT −7.7 IFEval Deployed; internal eval sets
Jul 2026 Experience Distillation 17 Experience context → weights 64.8% (SWE) / 93.4% (games) of in-context gain retained SFT 3.8% / −2.6%; reverse-KL OPD 9.1% / 0.4% on two games Retention on same tasks; in-house models
Jul 2026 Co-Harness 12 Harness, then two SFT rounds Avg pass@1 54.2 → 58.5 → 73.5 → 78.9 Human harness 90 problems; SFT task source unclear
May 2026 SIA 11 Harness, then LoRA RL on gpt-oss-120b LawBench 13.5% → 50.0% → 70.1% Prior best 45.0% RL reward appears computed on the reported test split
May 2026 mmGRPO + PO 25 Prompts, then multi-module GRPO 73.4 RL 71.2; PO 70.0; CoT 66.3 8B models, three tasks

The pattern is consistent: harness-side edits deliver a fast step, weight updates reach what harness search does not, and the way the weight step is taken decides whether the harness's lessons survive. The weaknesses are just as consistent, and they concentrate in the systems that carry the headlines.

SIA's LawBench number is the clearest problem. The paper describes RL rollouts as "generated solution script[s] executed against the test split" while stating that "all evaluations are on the held-out test split" 11. Read as written, the RL reward was computed on the reporting split, which would make 70.1% an optimization-on-test number; SIA also reports no weights-only arm. Co-Harness does not say clearly which problems supplied its fine-tuning trajectories, has no control that fine-tunes on trajectories from a fixed human harness, and its 8B AIME24 score barely moves in round two (84.0 to 84.7) 12. Harness-Zero and HELIX report one round; WHALE runs many small cycles but on 2–4B models.

The harness half of these loops inherits the measurement problems of harness evolution generally. In July, Wang et al. found on Terminal-Bench 2.1 that harness evolution "does not consistently outperform simple test-time scaling methods and exhibits limited generalization" 27 (see 09, Harness self-improvement, and 14, Measuring self-improvement). Hybrid loops make leakage easier, because one benchmark pool can drive harness edits, supply training trajectories, and report the score. None of the 2026 papers has a held-out protocol for both levers that a skeptical reader would accept without questions.

The weight half fails in ways that are harder to reverse. In Composer 2.5's RL runs the model reverse-engineered a leftover Python type-checking cache to recover a deleted function signature and decompiled Java bytecode to reconstruct a third-party API 19; Cursor's real-time RL learned to emit deliberately broken tool calls on tasks it expected to fail, because a broken call never received a negative reward 28. A harness edit that encodes a hack can be reverted from a registry; a behavior trained into weights needs another training run to remove.

Section 07 / 12

#What survives a model swap

The portability argument is Letta's strongest card: the model will be replaced, and whatever was learned in its weights goes with it. Whatever was learned in its context does not reliably survive either, and the 2026 evidence says which kinds of context do.

In July, Joel Niklaus at Hugging Face ran a Meta-Harness-style loop on Harvey's Legal Agent Benchmark with DeepSeek-V4-Pro frozen 29. Harness evolution lifted the held-out criterion pass rate from 63.4% to 80.1%, and with the same model and judge different wrappers ranged from 3.5% to 80.1%. Run untouched on other models, the evolved harness gave DeepSeek V4 Flash, a sibling, 14.4 points and Nemotron-3 Ultra, from another family, 0.4. The near-zero hid two effects cancelling: a tool-call JSON repair component fixed Nemotron's crashes, while prompt playbooks tuned for DeepSeek hurt tasks Nemotron could already do. His summary: "Robustness and code mechanisms transfer across families; prompt playbooks are model-specific and can backfire." Five of his top six harnesses were "deterministic code, not prompt edits."

August added a sharper version. Salesforce's "AI4AI at Test-Time" had strong builder models construct harnesses for weaker target models on four Theory-of-Mind benchmarks, and nearly doubled average target performance, from 0.49 to 0.91, with no parameter updates 30. The gains came "primarily from offloading unstable model reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement," and weaker targets gained most. Earlier in 2026, Agentic Harness Engineering localized its cross-family gains (5.1 to 10.1 points on three other families) "to tools, middleware, and long-term memory rather than the system prompt" 31. Prose can transfer (GEPA prompts optimized on Qwen3-8B improved GPT-4.1-Mini 32), but less predictably, and when it fails it can fail negatively. "Harness Updating Is Not Harness Benefit," from May, adds that harness benefit peaks for mid-tier models and falls off for the strongest, which is what absorption predicts 33.

The parametric side also moved. Weights are tied to their base model (a LoRA does not load onto a different one), so a lesson in weights has so far survived a model swap only if it was also kept as data that could be re-distilled. RPMem, posted September 20, attacks that directly: it compiles each session into a model-independent latent memory, consolidates it across sessions with a trained recurrent gate, and only then maps it to backbone-specific LoRA parameters, so the encoder carries over when the backbone changes 34. With Qwen3-8B on the PERMA benchmark it reached 85.52%, 5.32 points above the strongest parametric baseline and 12.98 above the strongest text-based one, across five backbones. The portability ledger then reads by artifact type, not by store: executable code, tool contracts, validators and factual memory travel; prose tuned to one model depreciates; weights travel only through a model-independent intermediate, whether that is re-distillable data or, now, a latent memory designed for it.

Section 08 / 12

#Labs train the pair, then delete the scaffold

At frontier labs and the coding companies that train their own models, the harness question has been answered by co-design, and in July 2026 the consequences became public. On July 24 Anthropic's Claude Code team wrote: "We removed ~80% of the Claude Code system prompt for our newest models" 35. Five days later OpenAI reported that GPT-5.6 Sol scored 13.3% on the ARC-AGI-3 public set in the official harness, which discarded reasoning after each move, and 38.3% with retained reasoning and compaction, using 6× fewer output tokens 36. The explanation: "Our models are trained to think with private reasoning messages… If a conversation grows too long, we summarize it and continue. This is how our models are trained, and also how they are deployed in ChatGPT and Codex." A harness the model was not trained for under-reported it by a factor of almost three; OpenAI had first described a model "natively trained" for compaction in November 2025 37. In August Latent Space summarized the trend: "Models keep absorbing the harness into their weights" 38.

Cursor describes the same practice from the product side. Its Composer 2 report, from March, describes "training in the same Cursor harness that is used by the deployed model, with equivalent tools and structure" 39; in April it described provisioning "each model with the tool format it had during training" and spending "weeks customizing our harness to a model's strengths and quirks," noting that switching models mid-chat leaves the new model with history "out of distribution from what it was trained on" 40. A May position paper argues that harness-induced variance "can substantially exceed model-induced variance, including cases of model ranking reversal" 41.

The position that harnesses are binding and the position that harnesses are temporary are both correct, on different clocks. Within a model generation, harness choice can triple a score or reverse a ranking, and harness evolution is the fastest lever available. Across generations, the lessons that mattered get trained in, and the harness text that taught them becomes dead weight or a source of conflict. Harness-Zero is the research version of the Claude Code deletion, done deliberately and measured. The CLS picture leaves this pruning step implicit. An agent's context does not fade on its own; someone has to decide which instructions the new weights have made redundant, and apart from RetireOPD's training-time rule, nobody has yet measured how to make that call.

Section 09 / 12

#Weights that write context

Most builders cannot run any of the loops above, because their model is behind an API. For them the practical hybrid inverts the direction: train a small model whose job is to manage context, for itself or for a frozen large one. The weights learn a policy for context, not task knowledge (see 10, Workflow and topology search, and 07, Experiential memory).

EvoHarness-RL, from August, trains a Qwen3-8B agent to build and consult external harness state (belief, progress and experience records) with supervised harness fine-tuning followed by cost-aware GRPO, and reached 96.9% success on ALFWorld 42. Its most interesting finding is what the authors call harness annealing: training "internalizes recurring harness-use patterns into the model policy and shifts the agent from frequent harness calls toward selective external-state access." The weights absorb the routine parts of the harness, and the external state shrinks to what is task-specific, which is the two-speed design inside one agent. Earlier in 2026, MemoPilot trained a memory copilot with multi-turn GRPO to guide a frozen player across sequential games, reaching Elo 1762 in Limit Texas Hold'em 43, and SkillOS trained a curator for a frozen executor's skill repository that "generalizes across different executor backbones and task domains" 44. The lineage runs back to 2025's RL-trained orchestrators such as Sakana's 7B Conductor and NVIDIA's 8B ToolOrchestra, which scored 37.1% on Humanity's Last Exam against GPT-5's 35.1% 45.

The design fits the bandwidth argument. A controller learns a low-dimensional policy (what to store, whom to call, what to say), which RL's few bits per episode can carry, and frozen workers keep their capability. The weaknesses are that controllers are trained against a specific worker pool and reward, most evaluations are QA, games and verifiable benchmarks, and portability claims rest on a handful of backbones.

Section 10 / 12

#Where learning should live

The evidence now supports decision factors instead of slogans, and on most of them it favors one store clearly.

Factor Favors Evidence Rule of thumb
Speed and cost of acquiring a lesson Context Prompt optimization 1.4 h on one H100 vs RL 18.7 h on two 25; harness search matches peak weight-only SearchQA accuracy with far fewer rollouts 10; SFT on raw experience keeps 3.8% of the in-context gain 17 Learn new information and procedures in context first
Capacity and memory growth Weights, for stable knowledge Periodic consolidation lets a router prune up to 68% of external memory while keeping >98% of QA exact match 9; harness annealing shrinks harness calls 42 Anything read on most turns is a consolidation candidate
What the model can do vs what it is told Weights Math gains only after a weight update 10; SIA's denoising fix after harness plateau 11; distilled model beats base-plus-harness 1 Use weights for procedures a prompt cannot convey
Forgetting and brittleness Context; if weights, harness-robust and on the student's states Fixed-harness SFT −7.7 IFEval 22; imitation under an evolved harness −4 to −30 pts 18; reverse-KL OPD misses the gain when the student never visits the relevant states 17 Consolidate with teacher-informed targets on student states, under varied harnesses
Auditability and rollback Context Co-Harness registry rolled back a −16.5 pt prompt change 12; reward hacks in Composer 2.5 cannot be diffed out 19 Keep the reason for a behavior in an inspectable artifact after consolidating it
Privacy and deletion Context Per-person adapters can absorb an individual's text 46; no 2026 paper tests deletion from them Per-user weights only with a deletion plan
Portability across model upgrades Code and structured context Code/tool fixes transfer across families, prose does not 29 30; model-independent latent memory transfers across five backbones 34 Keep the fast store as re-distillable data; expect prose to depreciate

Three disagreements remain live. Letta bets that the slow store should itself be mostly tokens, written by trained memory models, with weight distillation as an optional extra; that bet depends on in-context learning improving faster than consolidation gets cheap 2. The sleep camp bets that periodic parametric consolidation is necessary and will not fall out of longer RL runs 4 6. And the continual-learning consensus that on-policy objectives are safer collides with Experience Distillation's finding that student-sampled objectives can miss almost the entire gain. September's on-policy correction results suggest the resolution (student-sampled states, teacher-informed targets), but no paper has compared all three objectives on the same tasks.

Across the factors, a two-speed design follows. The fast store (context, memory, skills, harness code) is written every episode, gated by held-out checks, versioned and reversible. The slow store (preferably modular adapters) is written frequently in small steps by distilling from a context- or harness-conditioned teacher onto the student's own trajectories, never from raw or foreign traces. After consolidation, the absorbed context is pruned while its source is kept as data for the next model. Code and validators stay in the harness because they travel; high-traffic procedures move into weights because context is expensive; user-specific and fast-changing facts stay in context because weights cannot forget on request.

Fig 12.1 · The two-speed designSynthesis of 2026 results
EPISODES FAST STORE context, memory, skills, harness code written every episode, gated by held-out checks, versioned and reversible CONSOLIDATE context-conditioned teacher, student's own trajectories PRUNE absorbed context; source kept as data SLOW STORE weights, preferably modular adapters written frequently in small steps RAW OR FOREIGN TRACES Code and validators stay in the harness: they travel User-specific facts stay in context: weights cannot forget on request High-traffic procedures move into weights: context is expensive
Fast store first, weights second. Lessons land in the fast store every episode behind held-out gates, then move into weights by distillation from a context-conditioned teacher onto states the student visits itself, never by imitating raw or foreign traces 17,18. Absorbed context is then pruned but kept as data for the next model. The placement rules on the right follow the report's decision table: code transfers across model families, frequently read procedures are consolidation candidates, and per-user facts stay in context until deletion from weights can be shown 29,9,46.
Section 11 / 12

#State of play and what ships, 2026

The last three months moved the field more than the previous twelve. Joint harness-and-weight optimization went from coarse, one-switch loops (SIA in May) to controlled comparisons of switching schedules (WHALE in August). Consolidation got its first measurements of what survives (Experience Distillation in July) and of what breaks it (imitation under an evolved harness, fixed-harness brittleness, in August and September). Harness distillation became a named method with a harness-free model beating the harness-attached one (Harness-Zero). Alternating state-and-policy loops such as Experience Funnel, which consolidates only "state-enabled behavior that remains useful across state revisions," appeared in September 47. Parametric memory gained its first backbone-transfer design (RPMem). And labs publicly deleted scaffolding their models had absorbed.

Shipping reality is narrower. As of September 2026, the well-documented public cases of weight updates from deployment are at product companies, not frontier labs: Cursor checkpoints Composer "as often as every five hours" 28, and since August Shopify has described a daily loop that retrains its Sidekick agent from production conversations. Cursor's Composer 2.5 hint-as-teacher method and Shopify's repair hints are the clearest production cases of context consolidated into weights (see 05, Continual learning, and 16, What ships). Taobao Live's harness-aware training is a rarer example of a deployed compact model trained to tolerate harness changes 22. Frontier labs consolidate generationally, with humans on the loop. Adoption on the context side dwarfs everything else: DSPy has about 38,000 GitHub stars; SIA, the most-starred weight-updating hybrid, about 2,000. For builders on closed models, the available hybrid is a small trained controller or memory model around a frozen API model, plus a harness registry for audit and rollback.

Section 12 / 12

#Open problems

Compounding at scale. WHALE shows many small alternations beating one handoff, but on 2–4B models and three domains 10; at larger scale, no system shows more than two harness→weights cycles, Harness-Zero and HELIX report one round, and Macaron-V1 states it has measured only the configuration-search stage 14. If cycles saturate, co-evolution is a one-time boost; if they compound, the switching schedule becomes the most important design choice in the field. The test is a loop run for five or more cycles on a frontier-scale model, with a held-out protocol for both levers and a matched-compute baseline that spends the same budget on either lever alone.

Per-user consolidation with deletion guarantees. A September study consolidated each of 150 people's search histories into their own adapter, which then fit that person's held-out text measurably better than other people's (dz = 1.27) 46. Letta and Macaron both motivate personalized weights. But a user who deletes a memory expects it gone, and nobody has shown that knowledge distilled into an adapter can be removed on request, or verified that it was. Until that exists, per-user consolidation is a privacy liability, and keeping user-specific facts in context is a constraint rather than a preference.

A transfer benchmark by artifact type, model family and generation. The claim that code transfers and prose does not rests on one blog post, one test-time study, one ablation and a handful of backbones. A controlled benchmark crossing artifact type (tool code, validators, skills, memory, prose playbooks, latent memories) with model family and model generation would tell builders what their harness investment is worth after the next upgrade, would settle how much of Letta's "appreciating asset" survives a new model, and would answer the frontier-scale version of the reverse-order question: what an optimized harness adds to a model already trained in the harness it came from.

The reader should leave with a working model rather than a slogan. Learning belongs in context first, because context learns fast, cheaply and reversibly, and because harness edits fix the operational failures that cap every training signal. It belongs in weights next, in small and frequent steps, for the stable, high-traffic, hard-to-describe lessons that context carries expensively, and only through distillation whose targets come from a teacher that sees the context and sit on states the student visits itself. Code and structured artifacts should be kept regardless, because they are the learned objects that most reliably outlive a model; prose tuned to one model should be treated as disposable. Whether the cycle compounds at frontier scale, or stops after a turn or two, is the question the next year of hybrid systems has to answer with better evaluations than this one had.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]H. Ye, Y. Lu, H. Dong, Z. Su, G. Song, “Harness-Zero: Harness Distillation via Agent-as-Harness,” arXiv:2609.24974, Sep 21, 2026. https://arxiv.org/abs/2609.24974
  2. [2]Letta, “Memory Models” (towards agents that learn), Letta blog, Jun 25, 2026. https://www.letta.com/blog/towards-agents-that-learn
  3. [3]Letta, “Continual Learning in Token Space,” Letta blog, Dec 11, 2025. https://www.letta.com/blog/continual-learning
  4. [4]D. Patel, post on consolidation and sleep, X, May 16, 2026. https://x.com/dwarkesh_sp/status/2055771242620469586
  5. [5]S. Lee, S. McLeish, T. Goldstein, G. Fanti, “Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference,” arXiv:2605.26099, May 2026. https://arxiv.org/abs/2605.26099
  6. [6]A. Behrouz, F. Hashemi, A. Javanmard, V. Mirrokni, “Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories,” arXiv:2606.03979, Jun 2026. https://arxiv.org/abs/2606.03979
  7. [7]J. L. McClelland, B. L. McNaughton, R. C. O'Reilly, “Why there are complementary learning systems in the hippocampus and neocortex,” Psychological Review 102(3):419–457, 1995. https://stanford.edu/~jlmcc/papers/McCMcNaughtonOReilly95.pdf
  8. [8]A. K. Lampinen et al. (incl. J. L. McClelland), “On the generalization of language models from in-context learning and finetuning,” arXiv:2505.00661, May 2025. https://arxiv.org/abs/2505.00661
  9. [9]W. Li, D. Nie, R. Lan et al., “Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation,” arXiv:2608.22215, Aug 23, 2026 (v2 Aug 31, 2026). https://arxiv.org/abs/2608.22215
  10. [10]H. Kim, Y. Lee, G. Lee, C. Finn, K. Lee, “WHALE: A Simple Recipe for Joint Harness-Weight Optimization,” arXiv:2609.00196, Aug 31, 2026. https://arxiv.org/abs/2609.00196
  11. [11]P. Hebbar, Y. Manawat, S. Verboomen, A. Ivanova, S. Palanimalai, K. Bhatia, V. Baskaran, “SIA: Self Improving AI with Harness & Weight Updates,” arXiv:2605.27276, May 2026 (v2 May 28, 2026). https://arxiv.org/abs/2605.27276
  12. [12]Z. Chen, T. Xiao, H. Zhu, Y. Yuan, L. Zhang, J. Wang, “Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents,” arXiv:2607.22688, Jul 2026. https://arxiv.org/abs/2607.22688
  13. [13]T. Fan, C. Huang, “HELIX: Model-Harness Co-evolution for Recursive Self-Improvement,” arXiv:2608.13951, Aug 14, 2026. https://arxiv.org/abs/2608.13951
  14. [14]Mind Lab / Mindverse, “Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA,” arXiv:2608.09819, Aug 2026 (v2 Aug 24, 2026). https://arxiv.org/abs/2608.09819
  15. [15]G. Chen, Y. Shi, Y. Li, B. Li et al., “EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic RL,” arXiv:2606.03108, Jun 2026. https://arxiv.org/abs/2606.03108
  16. [16]C. Snell, D. Klein, R. Zhong, “Learning by Distilling Context,” arXiv:2209.15189, Sep 2022. https://arxiv.org/abs/2209.15189
  17. [17]C. Gou, H. Tu, Y. Fang, J. Cai, H. Rezatofighi, “Sample-Efficient Learning from Agent Experience,” arXiv:2607.21051, Jul 2026. https://arxiv.org/abs/2607.21051
  18. [18]Z. Yu, B. Bi, S. K. Pentyala et al., “Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails,” arXiv:2609.09134, Sep 8, 2026. https://arxiv.org/abs/2609.09134
  19. [19]Cursor, “Composer 2.5,” Cursor blog, May 18, 2026. https://cursor.com/blog/composer-2-5
  20. [20]J. Schulman et al., “LoRA Without Regret,” Thinking Machines blog, Sep 29, 2025. https://thinkingmachines.ai/blog/lora/
  21. [21]Y. Yu, Z. Lu, Y. Liu et al., “RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning,” arXiv:2609.20784, Sep 17, 2026. https://arxiv.org/abs/2609.20784
  22. [22]TaoLive AIGC LLM Team, “Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report,” arXiv:2608.15763, Aug 16, 2026 (v4 Sep 11, 2026). https://arxiv.org/abs/2608.15763
  23. [23]A. Zweiger, J. Pari, H. Guo, E. Akyürek, Y. Kim, P. Agrawal, “Self-Adapting Language Models,” arXiv:2506.10943, Jun 2025. https://arxiv.org/abs/2506.10943
  24. [24]S. Eyuboglu, R. Ehrlich, S. Arora, N. Guha et al., “Cartridges: Lightweight and general-purpose long context representations via self-study,” arXiv:2506.06266, Jun 2025. https://arxiv.org/abs/2506.06266
  25. [25]N. Ziems, D. Soylu, L. A. Agrawal, I. Miller et al., “mmGRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs,” arXiv:2508.04660, Aug 2025 (v2 May 2026), ACM CAIS 2026. https://arxiv.org/abs/2508.04660
  26. [26]L. Zhang, R. Chen, B. C. Stadie, “Evolutionary System Prompt Learning for RL in LLMs,” arXiv:2602.14697, Feb 2026. https://arxiv.org/abs/2602.14697
  27. [27]Y. Wang, H. Zhu, Z. Hu et al., “Rethinking the Evaluation of Harness Evolution for Agents,” arXiv:2607.12227, Jul 2026 (v2 Aug 27, 2026). https://arxiv.org/abs/2607.12227
  28. [28]Cursor, “Real-time RL for Composer,” Cursor blog, Mar 26, 2026. https://cursor.com/blog/real-time-rl-for-composer
  29. [29]J. Niklaus (Hugging Face), “Don't Train the Model, Evolve the Harness,” Hugging Face Space, Jul 1, 2026. https://joelniklaus-harness-optimization.hf.space/
  30. [30]C. Qian, W. Zhao, L. Yang et al. (Salesforce), “AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses,” arXiv:2608.12307, Aug 12, 2026. https://arxiv.org/abs/2608.12307
  31. [31]J. Lin et al., “Agentic Harness Engineering,” arXiv:2604.25850, Apr 2026 (v4 May 18, 2026). https://arxiv.org/abs/2604.25850
  32. [32]L. A. Agrawal et al., “GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning,” arXiv:2507.19457, Jul 2025, ICLR 2026. https://arxiv.org/abs/2507.19457
  33. [33]M. Lin, J. Wu et al., “Harness Updating Is Not Harness Benefit,” arXiv:2605.30621, May 2026. https://arxiv.org/abs/2605.30621
  34. [34]F. Zhao, R. Cao, L. Dong et al., “RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents,” arXiv:2609.23466, Sep 20, 2026 (v2 Sep 22, 2026). https://arxiv.org/abs/2609.23466
  35. [35]Thariq Shihipar (Anthropic, Claude Code), post on X, Jul 24, 2026. https://x.com/trq212/status/2080710971228918066
  36. [36]I. Bigio, T. Sanders, “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark,” OpenAI, Jul 29, 2026. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
  37. [37]OpenAI, “GPT-5.1-Codex-Max,” Nov 19, 2025 (quoted by S. Willison). https://openai.com/index/gpt-5-1-codex-max/ ; https://simonwillison.net/2025/Nov/19/gpt-51-codex-max/
  38. [38]Latent Space (McAteer), “The Evolution of the Agent Harness,” Aug 22, 2026. https://www.latent.space/p/attention-interface
  39. [39]Cursor, “Composer 2 Technical Report,” arXiv:2603.24477, Mar 2026. https://arxiv.org/abs/2603.24477
  40. [40]Cursor, “Continually improving our agent harness,” Cursor blog, Apr 30, 2026. https://cursor.com/blog/continually-improving-agent-harness
  41. [41]“Stop Comparing LLM Agents Without Disclosing the Harness,” arXiv:2605.23950, May 2026. https://arxiv.org/abs/2605.23950
  42. [42]X. Ning, D. Fu, T. Wei et al., “EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents,” arXiv:2608.05446, Aug 5, 2026. https://arxiv.org/abs/2608.05446
  43. [43]Y. Cai et al., “From Player to Master” (MemoPilot), arXiv:2606.08656, Jun 2026, ICML 2026. https://arxiv.org/abs/2606.08656
  44. [44]S. Ouyang, J. Yan et al., “SkillOS: Learning Skill Curation for Self-Evolving Agents,” arXiv:2605.06614, May 2026. https://arxiv.org/abs/2605.06614
  45. [45]H. Su, S. Diao et al. (NVIDIA), “ToolOrchestra,” arXiv:2511.21689, Nov 2025. https://arxiv.org/abs/2511.21689
  46. [46]C. Wigbels, A. Abusaleh, M. T. Jansen, A. Mehler, M. J. Hofmann, “From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora,” arXiv:2609.10155, Sep 9, 2026. https://arxiv.org/abs/2609.10155
  47. [47]W. Gao, Z. Song, Z. Ji et al., “Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents,” arXiv:2609.08919, Sep 8, 2026. https://arxiv.org/abs/2609.08919