Research series · 2026 · Report 07 of 16
Report 07 of 16 · Part II · Non-parametric · Memory

Experiential memory

learning in token space

Writing experience into the context window is the most widely deployed form of agent self-improvement, and in 2026 it is also one of the most effective on the benchmarks its methods target: memory systems with verifier-gated writes and read-time curation add up to 16 to 18 points on agent benchmarks, and retrieval over past trajectories has been shown to solve a KL-constrained RL objective exactly. The general evidence is less kind. Self-graded memories inflate confident mistakes, plain in-context learning beats dedicated memory systems on a general continual-learning benchmark, gains swing with task order, and compaction silently deletes the rules that mattered. What gets stored, who grades it, and when it is pruned now matter more than how it is retrieved.

ContentsReport 07 · Non-parametric
$ tree ./07-experiential-memory
./07-experiential-memory
├── 01-a-memory-that-grades-itself# 355 words
├── 02-why-write-experience-into-context# 384 words
├── 03-how-the-loop-works# 1319 words · 1 figure
├── 04-what-the-evidence-shows# 386 words
├── 05-where-it-breaks# 721 words · 1 figure
├── 06-state-of-play-2026# 341 words
├── 07-what-ships# 266 words
└── 08-open-problems# 345 words
8 sections · 35 references · 2 figures
56.9% vs 54.0%
BIRD with an independent memory grader vs self-graded
53% → 10%
safety rules left after one vs five compactions
72.5%
best score on LongMemEval-V2 (RAG baseline 48.5%)
35
references · 18 from Jun 24 – Sep 24, 2026
17 min
reading time
Storage
Non-parametric: raw trajectories, insights, workflows, skills, playbooks, state records (edge cases: trained curators and memory managers, logit shifts at decode time)
Engines
LLM reflection and curation; similarity retrieval; read-time synthesis; non-parametric advantage estimates
Evaluator
Environment outcome or tests where available; verifier scores; self-judgment very often; human corrections in deployment
Loop timescale
Within a session to across deployments (weeks)
Loop closure
Closed loop in most research; human-on-the-loop with audit logs and rollback in products
Evidence maturity
Strong on home benchmarks; weak or negative on general continual-learning benchmarks; multi-seed, task-order, and matched-budget checks only arriving in August–September 2026

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 08

#A memory that grades itself

On June 29, 2026, Mohammad Asadolahi and four co-authors posted a paper with a plain title, "Memory Reward Inflation in Self-Improving LLM Agents," and a result that sums up the subfield 1. They took a text-to-SQL agent on BIRD and gave it three configurations. With no memory, it reached 52.4% execution accuracy. With a memory of past episodes that the agent scored itself, in the style of the case-based system Memento, it reached 54.0%. With the same memory de-inflated by their method, LUCID, which draws on a signal independent of the agent's own grading, it reached 56.9% 1. Self-graded memory helped. It also captured only about a third of the available gain, because the bias that made the agent write a wrong query was the same bias that made it rate that query highly, and retrieval then served the confident mistakes first. Stored experience is only as good as the signal that admits, ranks, and deletes it.

Earlier that month, on June 4, the authors of CL-Bench had published the general-purpose benchmark the memory literature had been missing: six expert-validated domains (software engineering, signal processing, outbreak forecasting, database querying, strategic games, demand forecasting) whose tasks share latent structure that an agent can learn online, scored with a gain metric that subtracts what the model already knew 2. Agents overfit to immediate observations or fail to reuse what they learned; CL-Bench found that dedicated memory systems do not fix the problem and that naive ICL outperforms systems dedicated to memory management 2. In August, a re-evaluation of two memory-based self-improving methods found that their improvement is "highly dependent on task order," and that the default task orders in prior papers act as "an implicit curriculum" 3.

Those results do not say experiential memory fails. Systems published in August and September 2026 report double-digit gains over strong memory baselines, several under matched budgets and multiple seeds. The results say where the gains come from: the grader of what gets written, the decision about what to keep in raw form, and the policy for what to delete. The retriever has become the least interesting part of the system.

Section 02 / 08

#Why write experience into context

Every self-improving system in this series answers the same questions: where the improvement is written, what proposes it, what decides it is better, who signs off, and whether the gain beats spending the same compute some other way. Experiential memory writes the improvement into non-parametric state (anything outside the weights): raw trajectories, distilled insights, workflows, skills, playbooks, or explicit records of current state. The proposer is usually the model reading its own traces. The evaluator varies more than anything else in the design, and it is where the failures concentrate.

Text wins deployment because it is cheap, legible, reversible, and works on closed models. An insight costs one inference call to write; an engineer can read it; deleting a line undoes it. Letta, the company that grew out of MemGPT, argued in December 2025 that "the weights are temporary; the learned context is what persists" 4, and sharpened the claim on June 25, 2026: "While models are a depreciating asset, learning and memory represented in token space is an appreciating asset" 5.

The lineage behind the 2026 systems is short. Reflexion (2023) carried verbal self-critiques from one attempt to the next 6; ExpeL (2023) extracted insights across tasks 7; Agent Workflow Memory (2024) stored reusable procedures and improved relative success by 51.1% on WebArena 8. Dynamic Cheatsheet (April 2025) let GPT-4o keep notes across a stream of problems and took it from about 10% to 99% on Game of 24, because the model wrote a Python solution once and reused it 9. ACE (October 2025) found the first structural failure: when an LLM rewrites a whole context each round, it compresses away detail (brevity bias) until the context collapses; in one run a playbook went from 18,282 tokens at 66.7 accuracy to 122 tokens at 57.1 in a single step, below the 63.7 of no adaptation at all 10. ACE's fix, itemized incremental updates with deliberate curation, became the default. ReasoningBank (ICLR 2026) distilled strategies from both successes and failures, but its labels came from an LLM judge rather than ground truth 11. Every stored lesson is only as good as that judge's call on which episodes went well.

Each of those systems decided at write time what an experience meant, and most let the model grade it. The 2026 work revisits both choices.

Section 03 / 08

#How the loop works

What to store: raw traces, curated at read time

The newest answer to "what should memory hold" is "less processing than you think." In September 2026, Zhou and colleagues argued that write-time curation, turning each finished trajectory into a fixed reflection, workflow, skill, or strategy, forces the system to decide what matters "before the future query is known, irreversibly discarding information" 12. Their Just-in-Time Memory (JitMem) keeps raw trajectories and defers curation until a new task arrives; a curator then synthesizes a compact payload for that specific task from the retrieved traces 12. Because the payload is used immediately, the curator can be trained from the current task's success rather than from a reward that arrives many tasks later. JitMem beat no-memory agents and heuristic and learned write-time memory methods, improving on the strongest baseline by 16.2 absolute success-rate points on ALFWorld, 16.3 on WebShop, and 3.9 on τ²-bench 12. The detail that matters most: "even an untrained curator is already competitive with or surpasses these baselines" 12. Read-time curation, not training, supplies most of the gain.

The finding lines up with the skeptical benchmarks. If the raw record plus a smart reader beats a pre-digested memory, then plain in-context learning beating dedicated memory systems on CL-Bench 2 is the limiting case of the same effect, and so is LongMemEval-V2's May 2026 result that an off-the-shelf coding agent searching stored trajectories reached 69.3% against 72.5% for the best purpose-built system and 48.5% for the strongest retrieval baseline 13. It also fits the September 2026 portability study below, where only agents that kept the raw source history could repair a damaged memory 14.

Raw traces have one blind spot: they record what was said, not what is currently true. In August 2026, StateMemBench tested 234 multi-session scenarios in which facts, constraints, and decisions are revised, and graded whether answers reflected the current state or a superseded one; existing memory systems, retrieval baselines, and long-context baselines all struggled 15. A wrapper that tracks supersession explicitly lifted current-state accuracy by 32 to 67 points across six memory and retrieval backends, and a length- and cost-matched control attributed 15 to 32 of those points to the state structure rather than to extra context 15. The emerging design stores raw history for evidence and a small, explicitly versioned record of current state for decisions.

Who grades it: verifiers that persist

The Memory Reward Inflation paper calls this shortfall the Echo Gap and proves that no procedure can correct inflated memory scores without a signal whose errors are independent of the agent's own 1. August 2026 brought the first systems built around that constraint.

MemGuard (August 22, 2026) names two failure modes: "unreliable admission," where failed trajectories, accidental successes, and misleading observations enter memory because they look relevant, and "memory drift," where long-running banks accumulate duplicate, stale, and conflicting records 16. Its move is to treat verifier output "not as a one-shot filter, but as persistent lifecycle metadata": each candidate memory carries reward, confidence, label, and uncertainty descriptors that are reused at retrieval, conflict resolution, summarization, and archival 16. Evaluated on Terminal-Bench 2.0, SWE-Bench Verified, WebArena, and Mind2Web across four backbones, against four memory baselines and a verifier-only control under matched runtime budgets and averaged over five seeds, MemGuard had the best success rate and fewest steps in all 16 backbone-benchmark settings 16. Its margin over ReasoningBank, the strongest baseline, peaked at 7.9 success-rate points on WebArena and ran 2.4 to 3.5 points on the terminal and software-engineering benchmarks 16. The matched-budget and multi-seed protocol is as notable as the result; few earlier memory papers reported either.

Recuris (August 25, 2026) applies the same idea to skills. A working memory tracks task progress and selects skills from an experiential store, which turns each execution into evidence that localizes failures to specific memory components; a fixed meta-agent then makes "localized, validation-gated updates to Skill Memory" 17. Across four long-horizon benchmarks and ten models it improved success in 35 of 37 completed model-benchmark pairs, adding 17.8 points to GPT-5.6 Sol and 15.6 to Claude Opus 5 on τ-bench (taking Opus 5 to 87.9%), with the margin widening to 32.2 points on the longest tasks 17. The skill side of the loop belongs to report 08 (see 08, Skills and tools); the relevant point here is that the update is gated by validation, not by the model's opinion of its own run.

Deployment gives the clearest view of the grader's weight. "Learning on the Job" (The Memory Company, July 24, 2026) ran a frozen agent on τ-bench's banking domain and turned each episode's outcome into a natural-language rule 18. Against a static-retrieval control, a one-bit success or failure verdict lifted single-trial success to 1.6 times baseline; human corrections lifted it to 2.6 times, converting 22 of 84 tasks the baseline never solved 18. Same storage, same retrieval; the signal set the ceiling.

The theory: retrieval as reinforcement learning

JitRL (NUS, January 2026; ICML 2026 spotlight) explains why a well-graded store works at all. It retrieves past trajectories similar to the current state, estimates each candidate action's advantage from what followed it, and adds those advantages to the model's output logits before sampling 19. The authors prove that this additive update is the exact closed-form solution to the KL-constrained policy optimization objective, the objective that KL-regularized fine-tuning optimizes with gradient steps 19. For a policy that must stay close to its base model, then, retrieval plus a logit shift is the optimum, not an approximation. JitRL set a new state of the art among training-free methods on WebArena and Jericho and beat the fine-tuned WebRL agent at more than 30 times lower monetary cost 19.

The theorem sharpens the Echo Gap rather than escaping it: the update is only as good as the stored advantages, and inflated rewards produce an exactly optimal step in the wrong direction. It also carries an operational catch. JitRL needs logit access, which means open weights or an API that exposes per-token logit biases 19, so the most principled method in the family runs on the smallest share of deployed agents.

When to prune: compaction, eviction, invalidation

Stores that only grow degrade, and every long-running agent eventually summarizes or deletes. August and September 2026 measured what that costs.

The Compaction Cliff paper (August 24, 2026) tested Claude Code's /compact prompt on Sonnet 4.6 across 20 production agent configurations and found it preserved 53% of safety rules after one compaction round and 10% after five 20. The mechanism is ACE's brevity bias in production: a safety rule and an episodic log compete for the same tokens and are summarized at the same rate, but only the rule needs its exact wording to stay enforceable. The authors' fix, Knowledge Triage, classifies each line of an agent's knowledge base by type and gives each type its own retention policy; its compactor preserved 2 to 4 times more safety rules than the strongest single-shot LLM compactor at every compression ratio, with 96% recall over five rounds 20.

Eviction is worse, because it is often irreversible. A September 2026 audit on LongMemEval-S reinstated the gold evidence for each wrongly answered question and reran the reader. Under top-k retrieval at an 80k-token budget, 67% to 73% of the errors that restoration fixed were irreversible for FIFO, random, and redundancy-aware eviction (60% for LLM-judged importance); at an 8k budget, all of them were 21. Once evicted, the information was gone.

Fig 07.1 · What compaction and eviction destroyTwo studies · Aug–Sep 2026
COMPACTION · CLAUDE CODE /COMPACT SAFETY RULES KEPT (%) After 1 round 53 After 5 rounds 10 0 100 Knowledge Triage (per-type retention): 96% recall over five rounds; 2 to 4 times more safety rules than the strongest single-shot LLM compactor EVICTION · LONGMEMEVAL-S, TOP-K IRREVERSIBLE SHARE OF FIXED ERRORS (%) 80K-TOKEN BUDGET FIFO, random, redundancy-aware 67–73 LLM-judged importance 60 8K-TOKEN BUDGET All policies all 0 100
Pruning deletes what mattered. Left: Claude Code's /compact prompt on Sonnet 4.6, across 20 production agent configurations, kept 53% of safety rules after one round and 10% after five; the Knowledge Triage compactor reached 96% recall over five rounds, a different metric reported in the same paper 20. Right: a separate restore-counterfactual audit on LongMemEval-S, where the bars give the share of restoration-fixed errors that eviction had made irreversible 21.

Staleness is the third leg. Invalidation contracts (August 31, 2026) attach version stamps to cached fixes learned from API errors so a client can evict entries after server-side drift without trial and error 22. Stamp validity is deterministic, but whether the planner uses the surviving fixes depends on the model: the same bytes produced 100% first-try compliance on Claude Haiku 4.5 and 11% or less on Claude Sonnet 5 22. Pruning, like grading, now has model-specific behavior that has to be tested.

Section 04 / 08

#What the evidence shows

On their home benchmarks, the gains are large and cheap to buy.

Date System Stored unit / intervention Grader Benchmark and model Result Ref
2026-09 JitMem Raw trajectories, curated at read time Task success (curator training) ALFWorld, WebShop, τ²-bench +16.2, +16.3, +3.9 abs over strongest baseline 12
2026-09 Memory portability study Notes vs fixed-schema graph under writer swap Exact-match answer codes 48 synthetic histories; two <10B open models Notes shift +9.91 or −13.28 pts; graph +0.0004 14
2026-08 Recuris Validation-gated skill memory Validation τ-bench (GPT-5.6 Sol, Claude Opus 5) +17.8, +15.6 pts; 35 of 37 pairs improve 17
2026-08 Compaction Cliff Claude Code /compact (Sonnet 4.6) None 20 production configurations 53% of safety rules kept after 1 round, 10% after 5 20
2026-08 MemGuard Verifier metadata on every record Multi-criteria verifier WebArena, Mind2Web, TB 2.0, SWE-Bench Verified; 4 backbones, 5 seeds, matched runtime Best in 16/16 settings; up to +7.9 over ReasoningBank 16
2026-08 StateMem wrapper Explicit current-state records Closed-pool grading StateMemBench, 6 backends +32 to +67 pts; +15 to +32 under length/cost match 15
2026-08 Memory poisoning 1.2% false assertions injected n/a LongMemEval Accuracy 0.850 → 0.300 23
2026-07 Learning on the Job Natural-language rules Outcome bit; human correction τ-bench banking; Mistral Large, Claude Sonnet 5 1.6× baseline; 2.6× with corrections 18
2026-06 Memory Reward Inflation Graded episodes Self-grade vs LUCID BIRD text-to-SQL 52.4% none; 54.0% self-graded; 56.9% LUCID 1
2026-06 CL-Bench Dedicated memory vs naive ICL Gain metric 6 expert-validated domains Naive ICL beats memory systems 2
2026-05 LongMemEval-V2 Stored trajectories Curated questions 451 questions, up to 115M tokens Best 72.5%; generic coding agent 69.3%; RAG 48.5% 13
2026-01 JitRL Trajectories + advantages → logits Environment reward WebArena, Jericho SOTA training-free; beats WebRL at >30× lower cost 19

Three patterns run through the table. The largest gains come where something other than the agent grades what gets stored: validation for Recuris, a verifier for MemGuard, task success for JitMem's curator, human corrections for Learning on the Job. Evaluation hygiene improved sharply in August: MemGuard ran five seeds under matched runtime, StateMem controlled for length and cost, and the fragility study showed why that matters, since stacking a self-improving loop on a noisy agent evaluation "can further amplify this noise" 3. And the general benchmarks still disagree with the home benchmarks: none of the systems at the top of the table has yet been run on CL-Bench.

Section 05 / 08

#Where it breaks

The Echo Gap

The mechanism is simple once stated 1. The bias that caused a wrong answer inflates its self-assigned score. Every retrieval of a stored episode acts as a policy-improvement step whose reliability depends on that score, so inflated episodes steer future behavior toward past mistakes. A second LLM judge does not help, because its errors are correlated with the first grader's. The shortfall compounds whether memories are ranked by score or by similarity. The paper's central result is that an Error-Independence Assumption (the correcting signal's errors must be independent of the errors that produced the episode) is necessary, not only sufficient, for any correction 1. LUCID works on BIRD because it has such a signal; in a deployment without one, the theorem says no algorithm recovers the loss. Report 03 treats the general version, where self-generated rewards become the objective (see 03, Self-generated rewards).

Fig 07.1 · The Echo Gap in a self-graded memoryBIRD text-to-SQL · June 2026
ONE BIAS WRITES AND GRADES LUCID: independent signal de-inflates scores Agent answers writes a wrong query Agent grades itself rates that query highly Episode stored with an inflated score Retrieved next time confident errors first each retrieval steers future behavior toward past mistakes BIRD TEXT-TO-SQL · EXECUTION ACCURACY (%) No memory 52.4 Self-graded memory 54.0 +1.6 OVER NO MEMORY LUCID (de-inflated) 56.9 +2.9 OVER SELF-GRADED 0 100
Self-grading captures about a third of the gain. The bias that produces a wrong query also grades it, so the stored episode carries an inflated score and is retrieved first on later tasks. On BIRD, self-graded memory in the style of Memento beat no memory by 1.6 points and left 2.9 points on the table relative to LUCID, which de-inflates scores with a signal independent of the agent's own grading 1.

The BIRD numbers keep the finding honest. Self-graded memory beat no memory by 1.6 points and left 2.9 on the table 1. The problem is inflation, not net harm, which is why products can ship self-graded memory and see gains while the store quietly drifts toward the model's blind spots.

Fragile gains

The August 2026 fragility study is the most direct challenge to the literature's headline numbers 3. Re-running two memory-based self-improving methods with multiple seeds and shuffled task orders, the authors found the loop amplifies evaluation noise and that improvement depends heavily on order, because default orderings encode a curriculum that the method silently relies on. Manual inspection of the memories pointed to task and environment underspecification; adding detailed rubrics and environment feedback to memory construction only partly closed the gap 3. A single-run, fixed-order memory result is now weak evidence.

General benchmarks disagree with home benchmarks

CL-Bench is not isolated. EvoMemBench (May 2026) evaluated 15 memory methods against long-context baselines under one protocol and found that "long-context baselines remain highly competitive, memory helps most when the current context is insufficient or tasks are difficult, and no single memory form works consistently across all settings" 24. LongMemEval-V2's generic coding agent nearly matched the best memory system 13. In September 2026, a team including Mem0's founders released DolphinBench, which scores memory through task completion on histories of about 500k tokens per persona, verifies that each task succeeds with the relevant history and fails without it, and requires cost and latency to be reported alongside accuracy 25. Benchmarks are converging on a standard that home-benchmark results rarely meet: beat long-context ICL, on tasks that require the history, at stated cost. Report 14 covers these benchmarks in depth (see 14, Measuring self-improvement).

Poisoned stores

Anything an agent writes, it later reads as instruction. MINJA (NeurIPS 2025) showed that query-only attackers can inject malicious records into a memory bank 26. An August 2026 study measured the cost with plainly worded false assertions: poisoning 1.2% of a LongMemEval corpus cut accuracy from 0.850 to 0.300, and a four-stage write-time screening pipeline that caught 83% of indirect prompt injections rejected 0 of 360 poisoned memories 23. The author's diagnosis matches the Echo Gap's: "distinguishing a false assertion from a true one generally requires external grounding beyond the text itself" 23. Report 15 owns these failures (see 15, Failure modes and safety).

Prose does not survive a model upgrade

The portability argument for token-space learning assumes a new model reads an old model's notes and benefits. A September 2026 controlled study tested that directly on 48 synthetic histories with two open-weight models under 10 billion parameters 14. A fixed-schema knowledge graph transferred almost perfectly when the writer model was swapped (accuracy change +0.0004). Model-written natural-language notes did not: accuracy moved by +9.91 or −13.28 percentage points depending on the direction of migration. Repairing the notes from the store alone failed to reach 90% recovery in all 48 cases, while keeping the raw source history allowed recovery in 34 of 48 for one direction 14. The result matches Hugging Face's July 2026 harness-transfer finding that "prompt playbooks are model-specific and can backfire" while code and structure transfer 27, and it qualifies Learning on the Job's report that its natural-language rules helped across two models 18. The small models and synthetic histories limit the study, but its direction is consistent with every other transfer result in the series (see 12, Consolidation and co-evolution).

Section 06 / 08

#State of play, 2026

The subfield spent 2023 to 2025 establishing that stored experience helps and how to write it without destroying it. From June to September 2026 the work moved to the parts of the loop that are not retrieval. A keyword sweep of arXiv turns up about 65 agent-memory preprints posted between early July and September 23, 2026; the strongest of them are about graders (MemGuard, Recuris), lifecycle and hygiene (compaction, eviction, invalidation, state tracking), and evaluation (fragility, DolphinBench, portability).

The theory is settled enough to guide design. JitRL shows retrieval-plus-advantage is RL under a KL constraint 19; the Echo Gap paper shows what any correction must assume 1. Together they locate the weak link in the stored rewards.

The token-space camp has conceded a role for weights. Letta's June 25, 2026 post proposed "memory models," trained with memory-native reinforcement learning to write and curate token-space memories that carry across model generations 5: to make context learning work, Letta proposes training the weights of the component that writes the context. MemoPilot (ICML 2026, June) does the same thing concretely, training a memory-update model with multi-turn GRPO while the player model stays frozen 28. JitMem's trained curator is a third instance 12.

Consolidation is arriving from the other side. Experience Distillation (July 23, 2026) found that distilling from an experience-conditioned teacher kept 64.8% of the in-context gain on software-engineering tasks, while supervised fine-tuning on the same experience kept 3.8% 29. Dual-Layer Agentic Memory (August 23, 2026), explicitly modeled on complementary learning systems, routes incoming information through a small-to-large model cascade that pruned up to 68% of redundant memory while retaining over 98% of the exhaustive baseline's exact-match accuracy, then consolidates high-value memories into the weights by fine-tuning 30. At the product level, Anthropic's Claude Code team said in July 2026 that it removed about 80% of Claude Code's system prompt for its newest models 31. The two stores are converging on a two-speed design, fast learning in context and slow consolidation into weights, and the open argument is scheduling (see 12, Consolidation and co-evolution).

Section 07 / 08

#What ships

Shipped memory has converged on files, audit logs, and rollback, which is the storage layer the research recommends, but graders are mostly left to the customer. Anthropic's memory for Claude Managed Agents, in public beta since April 23, 2026, "mounts directly onto a filesystem" so the agent uses its ordinary bash and code tools; every change is tracked "with a detailed audit log," and teams "can roll back to an earlier version or redact content from history" 32. Rakuten, quoted in the announcement, claims its agents deliver "97% fewer first-pass errors at 27% lower cost and 34% lower latency" 32; that is a customer statement, not a published evaluation. Letta moved the same way in March 2026, replacing specialized memory-editing tools with bash over git-backed files it calls MemFS 33. The filesystem-and-general-agent design is the product version of the research finding that raw records plus a capable reader beat bespoke memory machinery.

The loops with the strongest graders are the ones that put a test or a human in front of the write. Anthropic's June 2026 post on dynamic workflows in Claude Code describes improving CLAUDE.md by mining sessions and review comments for repeated corrections, clustering them, adversarially checking whether each candidate rule "would have prevented a real mistake," and distilling only survivors 34. Decagon's Autopilot, launched June 9, 2026, turns production-conversation signals into proposed agent updates, runs them against regression tests and the original conversation, and stages them for human review 35. Both are experiential memory with the Echo Gap engineered out. Report 16 covers memory companies and enterprise improvement loops (see 16, What ships).

Section 08 / 08

#Open problems

Correcting inflation without labels. The Echo Gap theorem says self-graded stores cannot be de-inflated without an error-independent signal 1, and MemGuard shows persistent verifier metadata beats strong baselines when a verifier exists 16. Most deployments have only cheap, partial signals: a follow-up message, whether an edit persisted, whether a ticket reopened. Nobody has shown which of these are independent enough to grade memories. The answer decides whether production memory can close its own loop or whether every durable lesson needs a test or a human signature, as the shipped loops already assume.

When to distill memory into weights. Experience Distillation shows consolidation can keep most of an in-context gain or almost none, depending on the objective 29; Dual-Layer Agentic Memory shows a working prune-then-consolidate pipeline on QA 30. Missing are a rule for when a lesson is stable enough to consolidate, a test that the weights absorbed it, and a way to delete the absorbed context without losing the audit trail. Closed-model users cannot consolidate at all and depend on the next model having learned what their memory taught the last one.

What survives shift. Home-benchmark gains come from repeated structure; CL-Bench and EvoMemBench ask for learning under heterogeneity, and memory loses 2,24. The September 2026 portability study says structured records transfer across writer models and prose notes do not 14, and the fragility study says task order alone can make or break a result 3. A controlled comparison of stored unit (raw trace, state record, skill, prose lesson) against distribution shift, task order, and model upgrade, at matched budget against long-context ICL, would settle more than another memory architecture.

Experiential memory is real self-improvement on the tasks it is built for, cheap enough to deploy everywhere, and in JitRL's form provably the optimal update for a KL-regularized policy. As of September 2026 it is also weaker than a long context on general continual-learning tasks, sensitive to task order, lossy under compaction, and bound by whatever grades what it stores. The next gains will come from the grader and the garbage collector, not the retriever.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]M. Asadolahi, A. Amini, S. Talebi, A. Farhadi, A. Zamanifar, “Memory Reward Inflation in Self-Improving LLM Agents,” arXiv:2608.00017, submitted June 29, 2026. https://arxiv.org/abs/2608.00017
  2. [2]“Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments,” arXiv:2606.05661, June 4, 2026. https://arxiv.org/abs/2606.05661
  3. [3]Q. Ye, Y. Li, Y. Pruksachatkun, J. Zhang, C.-S. Wu, “On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification,” arXiv:2608.18066, August 18, 2026. https://arxiv.org/abs/2608.18066
  4. [4]Letta, “Continual Learning in Token Space,” Letta blog, December 11, 2025. https://www.letta.com/blog/continual-learning
  5. [5]Letta, “Memory Models: Towards Agents That Learn,” Letta blog, June 25, 2026. https://www.letta.com/blog/towards-agents-that-learn
  6. [6]N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, S. Yao, “Reflexion: Language Agents with Verbal Reinforcement Learning,” NeurIPS 2023, arXiv:2303.11366, March 2023. https://arxiv.org/abs/2303.11366
  7. [7]A. Zhao, D. Huang, Q. Xu, M. Lin, Y.-J. Liu, G. Huang, “ExpeL: LLM Agents Are Experiential Learners,” AAAI 2024, arXiv:2308.10144, August 2023. https://arxiv.org/abs/2308.10144
  8. [8]Z. Z. Wang, J. Mao, D. Fried, G. Neubig, “Agent Workflow Memory,” arXiv:2409.07429, September 2024. https://arxiv.org/abs/2409.07429
  9. [9]M. Suzgun, M. Yuksekgonul, F. Bianchi, D. Jurafsky, J. Zou, “Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory,” arXiv:2504.07952, April 2025. https://arxiv.org/abs/2504.07952
  10. [10]Q. Zhang et al., “Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models,” arXiv:2510.04618, October 2025 (v3 March 2026). https://arxiv.org/abs/2510.04618
  11. [11]S. Ouyang et al., “ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory,” ICLR 2026, arXiv:2509.25140, September 2025. https://arxiv.org/abs/2509.25140
  12. [12]Y. Zhou, Y. Li, Z. L. Liu, S. Yavuz, S. Joty, “Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents,” arXiv:2609.27334, September 23, 2026. https://arxiv.org/abs/2609.27334
  13. [13]“LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues,” arXiv:2605.12493, May 2026. https://arxiv.org/abs/2605.12493
  14. [14]A. Goyal, J. Ray, “Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability,” arXiv:2609.05339, September 4, 2026. https://arxiv.org/abs/2609.05339
  15. [15]X. Fan, M. Liu, R. Yang, S. Ouyang, J. Han, “Can Agent Memory Systems Track Evolving State?,” arXiv:2608.19652, August 20, 2026. https://arxiv.org/abs/2608.19652
  16. [16]H. Wang, G. Dong, H. Liang, Z. Zhang, J. Luo, C. Liu, “MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance,” arXiv:2608.21867, August 22, 2026. https://arxiv.org/abs/2608.21867
  17. [17]Z. Yu, Y. Wu, Z. Yin, K. Chen, Z. Zhao, M. Wang, “Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses” (Recuris), arXiv:2608.24876, August 25, 2026. https://arxiv.org/abs/2608.24876
  18. [18]V. Tablan, S. Taylor, K. Bernhem (The Memory Company), “Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents,” arXiv:2607.22157, July 24, 2026. https://arxiv.org/abs/2607.22157
  19. [19]Y. Li, Z. Lin, A. Deng, X. Zhang, Y. He, S. Ji, T. Cao, B. Hooi, “Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates,” ICML 2026 (spotlight), arXiv:2601.18510, January 2026 (v3 June 2026). https://arxiv.org/abs/2601.18510 ; https://openreview.net/forum?id=pLvye0zHUC
  20. [20]S. Zerhoudi, J. Mitrovic, M. Granitzer, “The Compaction Cliff in Long-Running AI Agent Memory,” arXiv:2608.22752, August 24, 2026. https://arxiv.org/abs/2608.22752
  21. [21]C. Shen, “What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory,” arXiv:2609.08279, September 8, 2026. https://arxiv.org/abs/2609.08279
  22. [22]M. Wu, A. Canedo, “Invalidation Contracts for Cross-Episode Agent Memory,” arXiv:2609.00243, August 31, 2026. https://arxiv.org/abs/2609.00243
  23. [23]A. Karunanidhi, “Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking,” arXiv:2608.21230, August 21, 2026. https://arxiv.org/abs/2608.21230
  24. [24]Y. Wang et al., “EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective,” arXiv:2605.18421, May 18, 2026 (v2 June 2026). https://arxiv.org/abs/2605.18421
  25. [25]S. Rathi, D. Yadav, T. Singh, “DolphinBench: Mapping the Pareto Frontier of Agent Memory,” arXiv:2609.24971, September 21, 2026. https://arxiv.org/abs/2609.24971
  26. [26]“Memory Injection Attacks on LLM Agents via Query-Only Interaction” (MINJA), NeurIPS 2025, arXiv:2503.03704. https://arxiv.org/abs/2503.03704
  27. [27]J. Niklaus (Hugging Face), “Don't Train the Model, Evolve the Harness,” Hugging Face Space, July 1, 2026. https://joelniklaus-harness-optimization.hf.space/
  28. [28]Y. Cai et al., “From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory” (MemoPilot), ICML 2026, arXiv:2606.08656, June 7, 2026. https://arxiv.org/abs/2606.08656
  29. [29]C. Gou, H. Tu, Y. Fang, J. Cai, H. Rezatofighi, “Sample-Efficient Learning from Agent Experience” (Experience Distillation), arXiv:2607.21051, July 23, 2026. https://arxiv.org/abs/2607.21051
  30. [30]W. Li, D. Nie, R. Lan, T. Lyu, P. Wang, L. Hong, “Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation,” arXiv:2608.22215, August 23, 2026. https://arxiv.org/abs/2608.22215
  31. [31]T. Shihipar (Anthropic, Claude Code), post on X, July 24, 2026. https://x.com/trq212/status/2080710971228918066
  32. [32]Anthropic, “Built-in memory for Claude Managed Agents,” Claude blog, April 23, 2026. https://claude.com/blog/claude-managed-agents-memory
  33. [33]Letta, “Letta's Next Phase,” Letta blog, March 16, 2026. https://www.letta.com/blog/our-next-phase/
  34. [34]T. Shihipar (Anthropic), “A harness for every task: dynamic workflows in Claude Code,” Claude blog, June 2, 2026. https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code
  35. [35]Decagon, “Autopilot,” Decagon blog, June 9, 2026. https://decagon.ai/blog/autopilot