#A memory that grades itself
On June 29, 2026, Mohammad Asadolahi and four co-authors posted a paper with a plain title, "Memory Reward Inflation in Self-Improving LLM Agents," and a result that sums up the subfield 1. They took a text-to-SQL agent on BIRD and gave it three configurations. With no memory, it reached 52.4% execution accuracy. With a memory of past episodes that the agent scored itself, in the style of the case-based system Memento, it reached 54.0%. With the same memory de-inflated by their method, LUCID, which draws on a signal independent of the agent's own grading, it reached 56.9% 1. Self-graded memory helped. It also captured only about a third of the available gain, because the bias that made the agent write a wrong query was the same bias that made it rate that query highly, and retrieval then served the confident mistakes first. Stored experience is only as good as the signal that admits, ranks, and deletes it.
Earlier that month, on June 4, the authors of CL-Bench had published the general-purpose benchmark the memory literature had been missing: six expert-validated domains (software engineering, signal processing, outbreak forecasting, database querying, strategic games, demand forecasting) whose tasks share latent structure that an agent can learn online, scored with a gain metric that subtracts what the model already knew 2. Agents overfit to immediate observations or fail to reuse what they learned; CL-Bench found that dedicated memory systems do not fix the problem and that naive ICL outperforms systems dedicated to memory management 2. In August, a re-evaluation of two memory-based self-improving methods found that their improvement is "highly dependent on task order," and that the default task orders in prior papers act as "an implicit curriculum" 3.
Those results do not say experiential memory fails. Systems published in August and September 2026 report double-digit gains over strong memory baselines, several under matched budgets and multiple seeds. The results say where the gains come from: the grader of what gets written, the decision about what to keep in raw form, and the policy for what to delete. The retriever has become the least interesting part of the system.
#Why write experience into context
Every self-improving system in this series answers the same questions: where the improvement is written, what proposes it, what decides it is better, who signs off, and whether the gain beats spending the same compute some other way. Experiential memory writes the improvement into non-parametric state (anything outside the weights): raw trajectories, distilled insights, workflows, skills, playbooks, or explicit records of current state. The proposer is usually the model reading its own traces. The evaluator varies more than anything else in the design, and it is where the failures concentrate.
Text wins deployment because it is cheap, legible, reversible, and works on closed models. An insight costs one inference call to write; an engineer can read it; deleting a line undoes it. Letta, the company that grew out of MemGPT, argued in December 2025 that "the weights are temporary; the learned context is what persists" 4, and sharpened the claim on June 25, 2026: "While models are a depreciating asset, learning and memory represented in token space is an appreciating asset" 5.
The lineage behind the 2026 systems is short. Reflexion (2023) carried verbal self-critiques from one attempt to the next 6; ExpeL (2023) extracted insights across tasks 7; Agent Workflow Memory (2024) stored reusable procedures and improved relative success by 51.1% on WebArena 8. Dynamic Cheatsheet (April 2025) let GPT-4o keep notes across a stream of problems and took it from about 10% to 99% on Game of 24, because the model wrote a Python solution once and reused it 9. ACE (October 2025) found the first structural failure: when an LLM rewrites a whole context each round, it compresses away detail (brevity bias) until the context collapses; in one run a playbook went from 18,282 tokens at 66.7 accuracy to 122 tokens at 57.1 in a single step, below the 63.7 of no adaptation at all 10. ACE's fix, itemized incremental updates with deliberate curation, became the default. ReasoningBank (ICLR 2026) distilled strategies from both successes and failures, but its labels came from an LLM judge rather than ground truth 11. Every stored lesson is only as good as that judge's call on which episodes went well.
Each of those systems decided at write time what an experience meant, and most let the model grade it. The 2026 work revisits both choices.
#How the loop works
What to store: raw traces, curated at read time
The newest answer to "what should memory hold" is "less processing than you think." In September 2026, Zhou and colleagues argued that write-time curation, turning each finished trajectory into a fixed reflection, workflow, skill, or strategy, forces the system to decide what matters "before the future query is known, irreversibly discarding information" 12. Their Just-in-Time Memory (JitMem) keeps raw trajectories and defers curation until a new task arrives; a curator then synthesizes a compact payload for that specific task from the retrieved traces 12. Because the payload is used immediately, the curator can be trained from the current task's success rather than from a reward that arrives many tasks later. JitMem beat no-memory agents and heuristic and learned write-time memory methods, improving on the strongest baseline by 16.2 absolute success-rate points on ALFWorld, 16.3 on WebShop, and 3.9 on τ²-bench 12. The detail that matters most: "even an untrained curator is already competitive with or surpasses these baselines" 12. Read-time curation, not training, supplies most of the gain.
The finding lines up with the skeptical benchmarks. If the raw record plus a smart reader beats a pre-digested memory, then plain in-context learning beating dedicated memory systems on CL-Bench 2 is the limiting case of the same effect, and so is LongMemEval-V2's May 2026 result that an off-the-shelf coding agent searching stored trajectories reached 69.3% against 72.5% for the best purpose-built system and 48.5% for the strongest retrieval baseline 13. It also fits the September 2026 portability study below, where only agents that kept the raw source history could repair a damaged memory 14.
Raw traces have one blind spot: they record what was said, not what is currently true. In August 2026, StateMemBench tested 234 multi-session scenarios in which facts, constraints, and decisions are revised, and graded whether answers reflected the current state or a superseded one; existing memory systems, retrieval baselines, and long-context baselines all struggled 15. A wrapper that tracks supersession explicitly lifted current-state accuracy by 32 to 67 points across six memory and retrieval backends, and a length- and cost-matched control attributed 15 to 32 of those points to the state structure rather than to extra context 15. The emerging design stores raw history for evidence and a small, explicitly versioned record of current state for decisions.
Who grades it: verifiers that persist
The Memory Reward Inflation paper calls this shortfall the Echo Gap and proves that no procedure can correct inflated memory scores without a signal whose errors are independent of the agent's own 1. August 2026 brought the first systems built around that constraint.
MemGuard (August 22, 2026) names two failure modes: "unreliable admission," where failed trajectories, accidental successes, and misleading observations enter memory because they look relevant, and "memory drift," where long-running banks accumulate duplicate, stale, and conflicting records 16. Its move is to treat verifier output "not as a one-shot filter, but as persistent lifecycle metadata": each candidate memory carries reward, confidence, label, and uncertainty descriptors that are reused at retrieval, conflict resolution, summarization, and archival 16. Evaluated on Terminal-Bench 2.0, SWE-Bench Verified, WebArena, and Mind2Web across four backbones, against four memory baselines and a verifier-only control under matched runtime budgets and averaged over five seeds, MemGuard had the best success rate and fewest steps in all 16 backbone-benchmark settings 16. Its margin over ReasoningBank, the strongest baseline, peaked at 7.9 success-rate points on WebArena and ran 2.4 to 3.5 points on the terminal and software-engineering benchmarks 16. The matched-budget and multi-seed protocol is as notable as the result; few earlier memory papers reported either.
Recuris (August 25, 2026) applies the same idea to skills. A working memory tracks task progress and selects skills from an experiential store, which turns each execution into evidence that localizes failures to specific memory components; a fixed meta-agent then makes "localized, validation-gated updates to Skill Memory" 17. Across four long-horizon benchmarks and ten models it improved success in 35 of 37 completed model-benchmark pairs, adding 17.8 points to GPT-5.6 Sol and 15.6 to Claude Opus 5 on τ-bench (taking Opus 5 to 87.9%), with the margin widening to 32.2 points on the longest tasks 17. The skill side of the loop belongs to report 08 (see 08, Skills and tools); the relevant point here is that the update is gated by validation, not by the model's opinion of its own run.
Deployment gives the clearest view of the grader's weight. "Learning on the Job" (The Memory Company, July 24, 2026) ran a frozen agent on τ-bench's banking domain and turned each episode's outcome into a natural-language rule 18. Against a static-retrieval control, a one-bit success or failure verdict lifted single-trial success to 1.6 times baseline; human corrections lifted it to 2.6 times, converting 22 of 84 tasks the baseline never solved 18. Same storage, same retrieval; the signal set the ceiling.
The theory: retrieval as reinforcement learning
JitRL (NUS, January 2026; ICML 2026 spotlight) explains why a well-graded store works at all. It retrieves past trajectories similar to the current state, estimates each candidate action's advantage from what followed it, and adds those advantages to the model's output logits before sampling 19. The authors prove that this additive update is the exact closed-form solution to the KL-constrained policy optimization objective, the objective that KL-regularized fine-tuning optimizes with gradient steps 19. For a policy that must stay close to its base model, then, retrieval plus a logit shift is the optimum, not an approximation. JitRL set a new state of the art among training-free methods on WebArena and Jericho and beat the fine-tuned WebRL agent at more than 30 times lower monetary cost 19.
The theorem sharpens the Echo Gap rather than escaping it: the update is only as good as the stored advantages, and inflated rewards produce an exactly optimal step in the wrong direction. It also carries an operational catch. JitRL needs logit access, which means open weights or an API that exposes per-token logit biases 19, so the most principled method in the family runs on the smallest share of deployed agents.
When to prune: compaction, eviction, invalidation
Stores that only grow degrade, and every long-running agent eventually summarizes or deletes. August and September 2026 measured what that costs.
The Compaction Cliff paper (August 24, 2026) tested Claude Code's /compact prompt on Sonnet 4.6 across 20 production agent configurations and found it preserved 53% of safety rules after one compaction round and 10% after five 20. The mechanism is ACE's brevity bias in production: a safety rule and an episodic log compete for the same tokens and are summarized at the same rate, but only the rule needs its exact wording to stay enforceable. The authors' fix, Knowledge Triage, classifies each line of an agent's knowledge base by type and gives each type its own retention policy; its compactor preserved 2 to 4 times more safety rules than the strongest single-shot LLM compactor at every compression ratio, with 96% recall over five rounds 20.
Eviction is worse, because it is often irreversible. A September 2026 audit on LongMemEval-S reinstated the gold evidence for each wrongly answered question and reran the reader. Under top-k retrieval at an 80k-token budget, 67% to 73% of the errors that restoration fixed were irreversible for FIFO, random, and redundancy-aware eviction (60% for LLM-judged importance); at an 8k budget, all of them were 21. Once evicted, the information was gone.
Staleness is the third leg. Invalidation contracts (August 31, 2026) attach version stamps to cached fixes learned from API errors so a client can evict entries after server-side drift without trial and error 22. Stamp validity is deterministic, but whether the planner uses the surviving fixes depends on the model: the same bytes produced 100% first-try compliance on Claude Haiku 4.5 and 11% or less on Claude Sonnet 5 22. Pruning, like grading, now has model-specific behavior that has to be tested.
#What the evidence shows
On their home benchmarks, the gains are large and cheap to buy.
| Date | System | Stored unit / intervention | Grader | Benchmark and model | Result | Ref |
|---|---|---|---|---|---|---|
| 2026-09 | JitMem | Raw trajectories, curated at read time | Task success (curator training) | ALFWorld, WebShop, τ²-bench | +16.2, +16.3, +3.9 abs over strongest baseline | 12 |
| 2026-09 | Memory portability study | Notes vs fixed-schema graph under writer swap | Exact-match answer codes | 48 synthetic histories; two <10B open models | Notes shift +9.91 or −13.28 pts; graph +0.0004 | 14 |
| 2026-08 | Recuris | Validation-gated skill memory | Validation | τ-bench (GPT-5.6 Sol, Claude Opus 5) | +17.8, +15.6 pts; 35 of 37 pairs improve | 17 |
| 2026-08 | Compaction Cliff | Claude Code /compact (Sonnet 4.6) |
None | 20 production configurations | 53% of safety rules kept after 1 round, 10% after 5 | 20 |
| 2026-08 | MemGuard | Verifier metadata on every record | Multi-criteria verifier | WebArena, Mind2Web, TB 2.0, SWE-Bench Verified; 4 backbones, 5 seeds, matched runtime | Best in 16/16 settings; up to +7.9 over ReasoningBank | 16 |
| 2026-08 | StateMem wrapper | Explicit current-state records | Closed-pool grading | StateMemBench, 6 backends | +32 to +67 pts; +15 to +32 under length/cost match | 15 |
| 2026-08 | Memory poisoning | 1.2% false assertions injected | n/a | LongMemEval | Accuracy 0.850 → 0.300 | 23 |
| 2026-07 | Learning on the Job | Natural-language rules | Outcome bit; human correction | τ-bench banking; Mistral Large, Claude Sonnet 5 | 1.6× baseline; 2.6× with corrections | 18 |
| 2026-06 | Memory Reward Inflation | Graded episodes | Self-grade vs LUCID | BIRD text-to-SQL | 52.4% none; 54.0% self-graded; 56.9% LUCID | 1 |
| 2026-06 | CL-Bench | Dedicated memory vs naive ICL | Gain metric | 6 expert-validated domains | Naive ICL beats memory systems | 2 |
| 2026-05 | LongMemEval-V2 | Stored trajectories | Curated questions | 451 questions, up to 115M tokens | Best 72.5%; generic coding agent 69.3%; RAG 48.5% | 13 |
| 2026-01 | JitRL | Trajectories + advantages → logits | Environment reward | WebArena, Jericho | SOTA training-free; beats WebRL at >30× lower cost | 19 |
Three patterns run through the table. The largest gains come where something other than the agent grades what gets stored: validation for Recuris, a verifier for MemGuard, task success for JitMem's curator, human corrections for Learning on the Job. Evaluation hygiene improved sharply in August: MemGuard ran five seeds under matched runtime, StateMem controlled for length and cost, and the fragility study showed why that matters, since stacking a self-improving loop on a noisy agent evaluation "can further amplify this noise" 3. And the general benchmarks still disagree with the home benchmarks: none of the systems at the top of the table has yet been run on CL-Bench.
#Where it breaks
The Echo Gap
The mechanism is simple once stated 1. The bias that caused a wrong answer inflates its self-assigned score. Every retrieval of a stored episode acts as a policy-improvement step whose reliability depends on that score, so inflated episodes steer future behavior toward past mistakes. A second LLM judge does not help, because its errors are correlated with the first grader's. The shortfall compounds whether memories are ranked by score or by similarity. The paper's central result is that an Error-Independence Assumption (the correcting signal's errors must be independent of the errors that produced the episode) is necessary, not only sufficient, for any correction 1. LUCID works on BIRD because it has such a signal; in a deployment without one, the theorem says no algorithm recovers the loss. Report 03 treats the general version, where self-generated rewards become the objective (see 03, Self-generated rewards).
The BIRD numbers keep the finding honest. Self-graded memory beat no memory by 1.6 points and left 2.9 on the table 1. The problem is inflation, not net harm, which is why products can ship self-graded memory and see gains while the store quietly drifts toward the model's blind spots.
Fragile gains
The August 2026 fragility study is the most direct challenge to the literature's headline numbers 3. Re-running two memory-based self-improving methods with multiple seeds and shuffled task orders, the authors found the loop amplifies evaluation noise and that improvement depends heavily on order, because default orderings encode a curriculum that the method silently relies on. Manual inspection of the memories pointed to task and environment underspecification; adding detailed rubrics and environment feedback to memory construction only partly closed the gap 3. A single-run, fixed-order memory result is now weak evidence.
General benchmarks disagree with home benchmarks
CL-Bench is not isolated. EvoMemBench (May 2026) evaluated 15 memory methods against long-context baselines under one protocol and found that "long-context baselines remain highly competitive, memory helps most when the current context is insufficient or tasks are difficult, and no single memory form works consistently across all settings" 24. LongMemEval-V2's generic coding agent nearly matched the best memory system 13. In September 2026, a team including Mem0's founders released DolphinBench, which scores memory through task completion on histories of about 500k tokens per persona, verifies that each task succeeds with the relevant history and fails without it, and requires cost and latency to be reported alongside accuracy 25. Benchmarks are converging on a standard that home-benchmark results rarely meet: beat long-context ICL, on tasks that require the history, at stated cost. Report 14 covers these benchmarks in depth (see 14, Measuring self-improvement).
Poisoned stores
Anything an agent writes, it later reads as instruction. MINJA (NeurIPS 2025) showed that query-only attackers can inject malicious records into a memory bank 26. An August 2026 study measured the cost with plainly worded false assertions: poisoning 1.2% of a LongMemEval corpus cut accuracy from 0.850 to 0.300, and a four-stage write-time screening pipeline that caught 83% of indirect prompt injections rejected 0 of 360 poisoned memories 23. The author's diagnosis matches the Echo Gap's: "distinguishing a false assertion from a true one generally requires external grounding beyond the text itself" 23. Report 15 owns these failures (see 15, Failure modes and safety).
Prose does not survive a model upgrade
The portability argument for token-space learning assumes a new model reads an old model's notes and benefits. A September 2026 controlled study tested that directly on 48 synthetic histories with two open-weight models under 10 billion parameters 14. A fixed-schema knowledge graph transferred almost perfectly when the writer model was swapped (accuracy change +0.0004). Model-written natural-language notes did not: accuracy moved by +9.91 or −13.28 percentage points depending on the direction of migration. Repairing the notes from the store alone failed to reach 90% recovery in all 48 cases, while keeping the raw source history allowed recovery in 34 of 48 for one direction 14. The result matches Hugging Face's July 2026 harness-transfer finding that "prompt playbooks are model-specific and can backfire" while code and structure transfer 27, and it qualifies Learning on the Job's report that its natural-language rules helped across two models 18. The small models and synthetic histories limit the study, but its direction is consistent with every other transfer result in the series (see 12, Consolidation and co-evolution).
#State of play, 2026
The subfield spent 2023 to 2025 establishing that stored experience helps and how to write it without destroying it. From June to September 2026 the work moved to the parts of the loop that are not retrieval. A keyword sweep of arXiv turns up about 65 agent-memory preprints posted between early July and September 23, 2026; the strongest of them are about graders (MemGuard, Recuris), lifecycle and hygiene (compaction, eviction, invalidation, state tracking), and evaluation (fragility, DolphinBench, portability).
The theory is settled enough to guide design. JitRL shows retrieval-plus-advantage is RL under a KL constraint 19; the Echo Gap paper shows what any correction must assume 1. Together they locate the weak link in the stored rewards.
The token-space camp has conceded a role for weights. Letta's June 25, 2026 post proposed "memory models," trained with memory-native reinforcement learning to write and curate token-space memories that carry across model generations 5: to make context learning work, Letta proposes training the weights of the component that writes the context. MemoPilot (ICML 2026, June) does the same thing concretely, training a memory-update model with multi-turn GRPO while the player model stays frozen 28. JitMem's trained curator is a third instance 12.
Consolidation is arriving from the other side. Experience Distillation (July 23, 2026) found that distilling from an experience-conditioned teacher kept 64.8% of the in-context gain on software-engineering tasks, while supervised fine-tuning on the same experience kept 3.8% 29. Dual-Layer Agentic Memory (August 23, 2026), explicitly modeled on complementary learning systems, routes incoming information through a small-to-large model cascade that pruned up to 68% of redundant memory while retaining over 98% of the exhaustive baseline's exact-match accuracy, then consolidates high-value memories into the weights by fine-tuning 30. At the product level, Anthropic's Claude Code team said in July 2026 that it removed about 80% of Claude Code's system prompt for its newest models 31. The two stores are converging on a two-speed design, fast learning in context and slow consolidation into weights, and the open argument is scheduling (see 12, Consolidation and co-evolution).
#What ships
Shipped memory has converged on files, audit logs, and rollback, which is the storage layer the research recommends, but graders are mostly left to the customer. Anthropic's memory for Claude Managed Agents, in public beta since April 23, 2026, "mounts directly onto a filesystem" so the agent uses its ordinary bash and code tools; every change is tracked "with a detailed audit log," and teams "can roll back to an earlier version or redact content from history" 32. Rakuten, quoted in the announcement, claims its agents deliver "97% fewer first-pass errors at 27% lower cost and 34% lower latency" 32; that is a customer statement, not a published evaluation. Letta moved the same way in March 2026, replacing specialized memory-editing tools with bash over git-backed files it calls MemFS 33. The filesystem-and-general-agent design is the product version of the research finding that raw records plus a capable reader beat bespoke memory machinery.
The loops with the strongest graders are the ones that put a test or a human in front of the write. Anthropic's June 2026 post on dynamic workflows in Claude Code describes improving CLAUDE.md by mining sessions and review comments for repeated corrections, clustering them, adversarially checking whether each candidate rule "would have prevented a real mistake," and distilling only survivors 34. Decagon's Autopilot, launched June 9, 2026, turns production-conversation signals into proposed agent updates, runs them against regression tests and the original conversation, and stages them for human review 35. Both are experiential memory with the Echo Gap engineered out. Report 16 covers memory companies and enterprise improvement loops (see 16, What ships).
#Open problems
Correcting inflation without labels. The Echo Gap theorem says self-graded stores cannot be de-inflated without an error-independent signal 1, and MemGuard shows persistent verifier metadata beats strong baselines when a verifier exists 16. Most deployments have only cheap, partial signals: a follow-up message, whether an edit persisted, whether a ticket reopened. Nobody has shown which of these are independent enough to grade memories. The answer decides whether production memory can close its own loop or whether every durable lesson needs a test or a human signature, as the shipped loops already assume.
When to distill memory into weights. Experience Distillation shows consolidation can keep most of an in-context gain or almost none, depending on the objective 29; Dual-Layer Agentic Memory shows a working prune-then-consolidate pipeline on QA 30. Missing are a rule for when a lesson is stable enough to consolidate, a test that the weights absorbed it, and a way to delete the absorbed context without losing the audit trail. Closed-model users cannot consolidate at all and depend on the next model having learned what their memory taught the last one.
What survives shift. Home-benchmark gains come from repeated structure; CL-Bench and EvoMemBench ask for learning under heterogeneity, and memory loses 2,24. The September 2026 portability study says structured records transfer across writer models and prose notes do not 14, and the fragility study says task order alone can make or break a result 3. A controlled comparison of stored unit (raw trace, state record, skill, prose lesson) against distribution shift, task order, and model upgrade, at matched budget against long-context ICL, would settle more than another memory architecture.
Experiential memory is real self-improvement on the tasks it is built for, cheap enough to deploy everywhere, and in JitRL's form provably the optimal update for a KL-regularized policy. As of September 2026 it is also weaker than a long context on general continual-learning tasks, sensitive to task order, lossy under compaction, and bound by whatever grades what it stores. The next gains will come from the grader and the garbage collector, not the retriever.
#References
● marks sources dated June 24 to September 24, 2026.
- [1]M. Asadolahi, A. Amini, S. Talebi, A. Farhadi, A. Zamanifar, “Memory Reward Inflation in Self-Improving LLM Agents,” arXiv:2608.00017, submitted June 29, 2026. https://arxiv.org/abs/2608.00017
- [2]“Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments,” arXiv:2606.05661, June 4, 2026. https://arxiv.org/abs/2606.05661
- [3]Q. Ye, Y. Li, Y. Pruksachatkun, J. Zhang, C.-S. Wu, “On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification,” arXiv:2608.18066, August 18, 2026. https://arxiv.org/abs/2608.18066
- [4]Letta, “Continual Learning in Token Space,” Letta blog, December 11, 2025. https://www.letta.com/blog/continual-learning
- [5]Letta, “Memory Models: Towards Agents That Learn,” Letta blog, June 25, 2026. https://www.letta.com/blog/towards-agents-that-learn
- [6]N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, S. Yao, “Reflexion: Language Agents with Verbal Reinforcement Learning,” NeurIPS 2023, arXiv:2303.11366, March 2023. https://arxiv.org/abs/2303.11366
- [7]A. Zhao, D. Huang, Q. Xu, M. Lin, Y.-J. Liu, G. Huang, “ExpeL: LLM Agents Are Experiential Learners,” AAAI 2024, arXiv:2308.10144, August 2023. https://arxiv.org/abs/2308.10144
- [8]Z. Z. Wang, J. Mao, D. Fried, G. Neubig, “Agent Workflow Memory,” arXiv:2409.07429, September 2024. https://arxiv.org/abs/2409.07429
- [9]M. Suzgun, M. Yuksekgonul, F. Bianchi, D. Jurafsky, J. Zou, “Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory,” arXiv:2504.07952, April 2025. https://arxiv.org/abs/2504.07952
- [10]Q. Zhang et al., “Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models,” arXiv:2510.04618, October 2025 (v3 March 2026). https://arxiv.org/abs/2510.04618
- [11]S. Ouyang et al., “ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory,” ICLR 2026, arXiv:2509.25140, September 2025. https://arxiv.org/abs/2509.25140
- [12]Y. Zhou, Y. Li, Z. L. Liu, S. Yavuz, S. Joty, “Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents,” arXiv:2609.27334, September 23, 2026. https://arxiv.org/abs/2609.27334
- [13]“LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues,” arXiv:2605.12493, May 2026. https://arxiv.org/abs/2605.12493
- [14]A. Goyal, J. Ray, “Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability,” arXiv:2609.05339, September 4, 2026. https://arxiv.org/abs/2609.05339
- [15]X. Fan, M. Liu, R. Yang, S. Ouyang, J. Han, “Can Agent Memory Systems Track Evolving State?,” arXiv:2608.19652, August 20, 2026. https://arxiv.org/abs/2608.19652
- [16]H. Wang, G. Dong, H. Liang, Z. Zhang, J. Luo, C. Liu, “MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance,” arXiv:2608.21867, August 22, 2026. https://arxiv.org/abs/2608.21867
- [17]Z. Yu, Y. Wu, Z. Yin, K. Chen, Z. Zhao, M. Wang, “Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses” (Recuris), arXiv:2608.24876, August 25, 2026. https://arxiv.org/abs/2608.24876
- [18]V. Tablan, S. Taylor, K. Bernhem (The Memory Company), “Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents,” arXiv:2607.22157, July 24, 2026. https://arxiv.org/abs/2607.22157
- [19]Y. Li, Z. Lin, A. Deng, X. Zhang, Y. He, S. Ji, T. Cao, B. Hooi, “Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates,” ICML 2026 (spotlight), arXiv:2601.18510, January 2026 (v3 June 2026). https://arxiv.org/abs/2601.18510 ; https://openreview.net/forum?id=pLvye0zHUC
- [20]S. Zerhoudi, J. Mitrovic, M. Granitzer, “The Compaction Cliff in Long-Running AI Agent Memory,” arXiv:2608.22752, August 24, 2026. https://arxiv.org/abs/2608.22752
- [21]C. Shen, “What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory,” arXiv:2609.08279, September 8, 2026. https://arxiv.org/abs/2609.08279
- [22]M. Wu, A. Canedo, “Invalidation Contracts for Cross-Episode Agent Memory,” arXiv:2609.00243, August 31, 2026. https://arxiv.org/abs/2609.00243
- [23]A. Karunanidhi, “Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking,” arXiv:2608.21230, August 21, 2026. https://arxiv.org/abs/2608.21230
- [24]Y. Wang et al., “EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective,” arXiv:2605.18421, May 18, 2026 (v2 June 2026). https://arxiv.org/abs/2605.18421
- [25]S. Rathi, D. Yadav, T. Singh, “DolphinBench: Mapping the Pareto Frontier of Agent Memory,” arXiv:2609.24971, September 21, 2026. https://arxiv.org/abs/2609.24971
- [26]“Memory Injection Attacks on LLM Agents via Query-Only Interaction” (MINJA), NeurIPS 2025, arXiv:2503.03704. https://arxiv.org/abs/2503.03704
- [27]J. Niklaus (Hugging Face), “Don't Train the Model, Evolve the Harness,” Hugging Face Space, July 1, 2026. https://joelniklaus-harness-optimization.hf.space/
- [28]Y. Cai et al., “From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory” (MemoPilot), ICML 2026, arXiv:2606.08656, June 7, 2026. https://arxiv.org/abs/2606.08656
- [29]C. Gou, H. Tu, Y. Fang, J. Cai, H. Rezatofighi, “Sample-Efficient Learning from Agent Experience” (Experience Distillation), arXiv:2607.21051, July 23, 2026. https://arxiv.org/abs/2607.21051
- [30]W. Li, D. Nie, R. Lan, T. Lyu, P. Wang, L. Hong, “Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation,” arXiv:2608.22215, August 23, 2026. https://arxiv.org/abs/2608.22215
- [31]T. Shihipar (Anthropic, Claude Code), post on X, July 24, 2026. https://x.com/trq212/status/2080710971228918066
- [32]Anthropic, “Built-in memory for Claude Managed Agents,” Claude blog, April 23, 2026. https://claude.com/blog/claude-managed-agents-memory
- [33]Letta, “Letta's Next Phase,” Letta blog, March 16, 2026. https://www.letta.com/blog/our-next-phase/
- [34]T. Shihipar (Anthropic), “A harness for every task: dynamic workflows in Claude Code,” Claude blog, June 2, 2026. https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code
- [35]Decagon, “Autopilot,” Decagon blog, June 9, 2026. https://decagon.ai/blog/autopilot