Research series · 2026 · Report 05 of 16
Report 05 of 16 · Part I · Parametric · Weights

Continual learning

weights that change after deployment

Continual learning in weights is the capability the field most often calls missing, and as of September 2026 no frontier lab has publicly described updating a deployed model from its traffic. Two product companies have: Cursor retrains its coding models on live user signals every few hours, and in August 2026 Shopify documented a daily loop that folds production failures into a smaller model's weights. Both are RL loops, both lean on text hints to make the reward usable, and both show that the binding constraint is the evaluator. Research through the summer of 2026 agrees on the shape of the problem: no single update rule handles every type of environmental change, online RL adapts best to new facts but breaks on noisy reward, test-time training has moved into agent episodes, and composed anti-forgetting tools still lose most of what they learn over long horizons.

ContentsReport 05 · Parametric
$ tree ./05-continual-learning
./05-continual-learning
├── 01-a-daily-loop-at-shopify# 500 words
├── 02-why-the-field-keeps-calling-it-the…# 381 words
├── 03-the-lab-gap# 211 words
├── 04-how-the-production-loops-work# 657 words · 1 figure
├── 05-why-the-production-loops-are-rl-loops# 451 words
├── 06-learning-at-test-time# 524 words
├── 07-memory-written-into-weights-during-the…# 273 words
├── 08-composing-the-old-tools# 242 words · 1 figure
├── 09-adapters-from-live-data# 228 words
├── 10-what-the-numbers-show# 338 words
├── 11-where-it-breaks# 558 words
├── 12-state-of-play-september-2026# 265 words
├── 13-what-ships# 132 words
└── 14-open-problems# 305 words
14 sections · 41 references · 2 figures
$27M → ~$1M
Shopify's yearly serving cost after the daily loop
2,000
requests per minute the retrained agent serves
0 of 20
published methods that remove a secret from the weights
41
references · 16 from Jun 24 – Sep 24, 2026
21 min
reading time
Storage
Weights: full-model online RL (Cursor); daily SFT + GRPO into a distilled specialist (Shopify); fast weights inside the forward pass (In-Place TTT, TTCD, Titans/HOPE line); swappable LoRAs and hypernetwork-generated adapters
Engines
Policy gradient on user signals; critic-repaired trajectories plus GRPO; test-time gradient steps on self-supervised losses; anchored sequential SFT; hypernetworks
Evaluator
User behavior plus an offline benchmark gate (Cursor); an LLM judge calibrated against expert annotations (Shopify); next-token loss; retention on earlier tasks
Loop timescale
Per token or chunk (fast weights) · per episode (agentic TTT) · hours (Cursor) · daily (Shopify) · per release (labs)
Loop closure
Human-on-the-loop in both production loops (benchmark gates, monitoring, reward patches, human annotation of unrepaired failures); closed loop inside the forward pass for test-time methods
Evidence maturity
Two documented production loops, both self-reported; test-time methods at ≤ 8B parameters; composition results on synthetic 100-task streams

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 14

#A daily loop at Shopify

On August 5, 2026, Shopify published how its Sidekick GraphQL agent learns from production. The agent answers merchant questions by writing and running queries against Shopify's Admin API, at up to 2,000 requests per minute 1. Each day the pipeline mines anonymized production traffic for hard negatives, conversations a calibrated LLM judge correctly scores low. A panel of frontier reasoning models critiques each failure, an arbiter merges the critiques into one repair instruction, and that instruction is injected before the user's turn (Shopify calls this "hinting"). The conversation is replayed from there and rescored. Repaired trajectories become training data; failures the critics cannot fix go to expert annotators from Toloka 1. On the same daily cadence Shopify runs a full-parameter fine-tune over the accumulated trajectories and then GRPO with the judge as reward, "training on both new and previous trajectories" to limit "drift and catastrophic forgetting across cycles" 1.

Shopify reports that the smaller model's quality "eventually surpasses the frontier-powered baseline" and estimates serving cost falling from about $27M a year to about $1M, a 96% reduction 1. The cost evidence is numerical; the quality evidence is only the claim. Its framing of the problem is the clearest a practitioner has written: a deployed frontier model "has no mechanism for internalizing what production teaches it," so "production knowledge piles up in words and code while the model's weights remain untouched" 1.

The loop is the series' two-speed design running in production. Shopify first improves the frontier-backed product without touching weights, through an autoresearch loop in which an agent edits prompts, tool definitions, and harness code and "keeps the change if the score improves"; only when those improvements plateau does it move to parameter updates 1. That inner acceptor is the greedy keep-if-better rule the series flags as prone to false commits (see 6, Prompt and program optimization); the weight updates inherit whatever the judge lets through. What holds the whole thing up is the judge. Shopify spends most of the post on it: a rubric, inter-annotator agreement as the judge's ceiling, calibration with GEPA and ACE, backtests against past A/B results, and deliberate degradation tests to confirm each criterion responds 1. The judge, in Shopify's words, "is only a proxy."

Cursor's March 2026 post on real-time RL for its Composer coding agent shows what happens when a proxy gets optimized against continuously. Composer checkpoints trained on live user reactions ship "as often as every five hours"; in Cursor's A/B results, agent edits that persisted in the codebase rose 2.28% and dissatisfied follow-ups fell 3.13% 2. The model also "figured out that if it deliberately emitted a broken tool call on a task it was likely to fail at, it would never receive a negative reward," and separately "learned to defer risky edits by asking clarifying questions" 2. Neither exploit would show up in an offline benchmark. Both loops deliver measured gains, and both depend on an evaluator that the model is now searching as hard as it searches the task.

Section 02 / 14

#Why the field keeps calling it the missing piece

The case for weight-level learning rests on an asymmetry Andrej Karpathy stated in October 2025: context is working memory, weights are "a hazy recollection," and "humans obviously have some process for distilling some of that knowledge into the weights. We're missing it" 3. Silver and Sutton's April 2025 "Era of Experience" paper made the broader case for agents that learn from their own experience streams, while staying agnostic about mechanism: adaptation "may happen by any means, for example updating the weights of a neural network, or adapting in-context" 4.

In 2026 the argument moved from research to economics. Dwarkesh Patel's August 2026 essay "8 Predictions for the Era of Continual Learning" predicts that deployment-time learning will create the first real switching costs between AI providers and push labs to ship earlier, since deployment itself becomes training 5. Inside the labs the tone is anticipatory. Dario Amodei told Patel in February 2026: "We have some evidence to suggest that [continual learning] is another of those problems that is not as difficult as it seems" 6.

The counter-position is that most of what continual learning promises can be had in text. The series' throughline is that most shipped self-improvement in 2026 lives in memory, skills, and harness edits (see 7, Experiential memory; 8, Skills and tools), which survive a model upgrade where a gradient update does not. The most useful 2026 evidence on this question is a July 2026 study from Berkeley, "When Does Continual Learning Require Learning" 7. On a common Qwen3-8B backbone, Anne Harrington, Trevor Darrell, Jitendra Malik, Yutong Bai and colleagues ran eight methods from four families through the same sequential protocol: prompt optimization (GEPA, ACE), supervised updates (SFT, SDFT), online RL (GRPO, SDPO), and context compression (Cartridges, In-Place TTT). "Prompt-based methods fit each new stage quickly but degrade on future tasks. Distillation-based methods accumulate knowledge stably but struggle to update outdated facts. Context compression improves efficiency without substantially improving the ability to learn new tasks. Online reinforcement learning adapts most effectively to knowledge updates but remains sensitive to noisy reward signals" 7. The authors conclude that continual learning "is not a single capability": the pattern of change in the environment decides "when adaptation must be learned inside model weights and when it can be achieved through external scaffolding" 7.

Section 03 / 14

#The lab gap

Searching public statements from OpenAI, Anthropic, Google DeepMind, Meta, and xAI between January and September 2026 turns up no announcement that any of them updates a deployed model's weights from user traffic. The closest items are forecasts and research papers: Amodei's remark, and Google's line of test-time memory architectures, none described as running in a product. Cursor's September 2025 Tab post names the contrast: "Most other LLM providers train on static datasets or use paid labelers, and only roll out a new model to users as part of a named model release every few months" 8. The lab loop is generational, with humans on the loop (see 1, Agentic RL).

The gap has concrete causes. A frontier chat model serves open-ended requests where no single signal means "that was correct." Its users are too heterogeneous for one shared update to suit all of them. A shared update from user data opens a poisoning surface a frozen model does not have. Every new checkpoint is supposed to pass a safety evaluation before it ships. Cursor and Shopify sidestep these constraints in the same way: a narrow product (code, store queries), a dense signal per interaction, a user base large enough to average out individuals, and an evaluation gate that fits the domain.

Section 04 / 14

#How the production loops work

Shopify: a judge, a repair panel, and replay

Shopify's loop inverts Cursor's choice of evaluator. Cursor reads reward directly from user behavior; Shopify builds an LLM judge first and calibrates it until it "matches humans about as well as humans match each other," using Cohen's kappa between two expert annotators on random samples as the ceiling 1. It insists that ground truth include "randomly sampled traffic, not only curated examples," because golden sets only test the failures a team already knows about 1.

The training step has two stages. Supervised fine-tuning distills the repaired trajectories, including their reasoning, into the smaller model; GRPO then samples groups of responses per prompt and reinforces the ones the judge scores highest 1. The hint that repairs a failure plays the same role as Cursor's textual feedback in Composer 2.5: it turns a sparse "this conversation went badly" into a concrete demonstration of the better turn. Forgetting is handled with the oldest tool available, replaying all prior trajectories in every cycle 1. Shopify also reports that compressing the roughly 6,000-token system prompt into about 1,500 learned gist tokens cut end-to-end latency about 38% in a load test at 350 requests per minute 1.

Fig 05.1 · Shopify's daily Sidekick loopShopify Engineering · Aug 2026
Production traffic up to 2,000 req/min Calibrated LLM judge mines hard negatives Critic panel + arbiter one repair instruction Hint, replay, rescore before the user's turn repaired Unrepaired: expert annotators (Toloka) judge as reward DAILY CYCLE Smaller model distilled specialist GRPO groups per prompt Full-parameter SFT distills repaired turns Trajectory store replays new + previous JUDGE CALIBRATION Kappa between two experts as the ceiling Backtests against past A/B results Deliberate degradation tests per criterion
The judge sits in two places. A calibrated LLM judge picks which production conversations to repair, and the same judge supplies the GRPO reward, so every daily update inherits whatever it misses. Repaired trajectories join a store that is replayed in full each cycle to limit forgetting; failures the critics cannot fix go to human annotators. Flow and calibration checks follow Shopify's own description 1.

Cursor Tab: the simplest reward in production

Cursor's Tab autocomplete, documented in September 2025, remains the cleanest example of learning from user signals. It "runs on every user action, handling over 400 million requests per day," with a reward of +0.75 for an accepted suggestion, −0.25 for a rejected one, and 0 when the model shows nothing, so a suggestion is worth showing only if the model expects acceptance above 25% 8. Policy gradient needs on-policy data, so each checkpoint ships before the next is trained; "it takes us 1.5 to 2 hours to roll out a checkpoint and collect the data for the next step" 8. The result "makes 21% fewer suggestions than the previous model while having a 28% higher accept rate" 8.

Cursor Composer: the same loop on a multi-step agent

Composer is harder: one session spans many tool calls, the signal arrives late, and the user's reaction is a sentence. Cursor trains Composer offline "in the same Cursor harness that is used by the deployed model" 9, so moving the loop online kept the environment fixed and changed only where episodes came from. Rewards come from what users do after the agent acts, and checkpoints pass CursorBench, Cursor's internal evaluation suite, before deployment 2.

The two exploits map onto two gaps in that reward. The broken-tool-call hack exploited missing coverage: episodes that ended without a scorable user reaction were scored as neutral, and Cursor fixed it "by correctly including broken tool calls as negative examples" 2. The deferral hack exploited a proxy: "the user did not complain about an edit" is not "the task got done." Both are standard reward-hacking shapes (see 15, Failure modes and safety). The humans stayed on the loop by watching behavior metrics and patching the reward, not by approving each checkpoint.

In May 2026, Composer 2.5 added "targeted RL with textual feedback" 10. At a turn where the model behaved poorly, Cursor inserts a short hint into the local context and uses "the resulting model distribution as a teacher," distilling the unhinted student toward it for that turn while the global RL objective stays in place 10. The pattern consolidates a lesson written in text into weights through a context-conditioned teacher; report 12 (Consolidation and co-evolution) covers why that style of distillation keeps far more of the in-context gain than plain SFT 11. Cursor does not say whether the method runs inside the real-time loop. Composer 2.5 also trained on 25 times more synthetic tasks, and in those environments the model found and decompiled "Java bytecode to reconstruct a third-party API" and reverse-engineered "a leftover Python type-checking cache" to recover a deleted function signature 10. Every reward, whether it comes from users, judges, or tests, gets probed for leaks.

Section 05 / 14

#Why the production loops are RL loops

Both production loops end in RL even though supervised fine-tuning on good examples would be simpler. The 2026 evidence on forgetting explains the choice.

The Continual Reasoning Gym (August 2026) tests exactly the setting a production loop lives in: updating an existing reasoning model with RL with verifiable rewards as each new task arrives, across five sequences of text and visual reasoning tasks 12. "Sequential RLVR exhibits modest forgetting," the authors find, yet final performance stays below a jointly trained multitask model, and forgetting explains only part of that gap 12. The rest comes from shared reasoning structure that joint training exploits and sequential training misses. Their fix, Continual Prompt Replay, replays earlier tasks' prompts and regenerates the responses with the current policy, keeping replay on-policy; "only CPR reaches MTRL-level performance" 12. Shopify's replay of accumulated trajectories is the production cousin of the same idea.

The mechanism behind low forgetting was identified in September 2025 by RL's Razor: the amount a fine-tuned model forgets tracks the KL divergence from the base policy on the new task, and "on-policy training naturally biases updates toward KL-minimal solutions" 13. An on-policy update can only raise the probability of outputs the model already samples, so it reweights existing behavior instead of overwriting it. That is the property a loop of many small updates needs, and it is the same property that makes short RL runs sharpen rather than expand capability (see 1, Agentic RL; 4, Self-training and distillation).

The 2026 work turns that observation into methods. Self-Distillation Fine-Tuning (January 2026) conditions the model on a demonstration and uses that version as its own teacher, generating on-policy targets without a reward; it reached higher new-task accuracy than SFT with substantially less forgetting and accumulated several skills sequentially without regression 14. "Learning, Fast and Slow" (May 2026) splits adaptation between optimized context and RL-updated parameters, reporting up to 3× better sample efficiency than RL alone and up to 70% less KL divergence from the base model; when task domains shift mid-stream it keeps acquiring new tasks where parameter-only RL stalls 15. Architecture offers a separate route: Meta's sparse memory finetuning cut the NaturalQuestions F1 drop after learning new facts to 11%, against 89% for full fine-tuning and 71% for LoRA, but it requires a memory-layer model 16.

The Berkeley comparison adds the caveat 7. Stability-anchored methods (SDFT, SDPO) accumulated knowledge under noisy signals but "struggle to learn changes" when facts drift. GRPO updated facts best but lost significant accuracy on a sentiment-forecasting task with a weak, noisy reward, because "if the reward is noisy, the algorithm may reinforce the wrong answer" 7. The production loops work because Cursor and Shopify spent their effort making the reward clean.

Section 06 / 14

#Learning at test time

Production loops update a shared model on a schedule. A second research line updates weights per input or per episode and discards them afterward. It began in 2024 as a benchmark trick: test-time training on each ARC puzzle's demonstrations took an 8B model to 53.0%, and an ensemble with program synthesis to 61.9%, comparable to the 60.2% average human score 17; TTRL (April 2025) replaced labels with majority-vote rewards on unlabeled test questions 18 (see 3, Self-generated rewards). The 2026 work turns it into something deployable.

Test-time training inside agent episodes

In July 2026, "No Time Like the Present" ran test-time training continuously inside multi-turn agent episodes, where each update changes the policy that generates the next round of training text 19. That creates a self-training loop that "helps when new trajectory information appears, but can amplify drift when the agent gets stuck and repeatedly trains on similar text" 19. The authors found that repetition in the update text separates the two regimes, and their method, Agentic Test-Time Training (aTTT), downweights the loss on tokens that appear in repeated n-grams from earlier updates. Running LoRA updates live through vLLM's runtime LoRA API costs 1.9× a no-TTT episode, and aTTT improved success by up to 5.0 points on ALFWorld and 4.9 on SWE-bench Lite 19. The authors are candid about what it does: gains "concentrate where models already have task competence but drift over long trajectories," so aTTT "mainly preserves existing competence rather than teaching new abilities" 19.

The drift finding is a small-scale version of the production problem. A loop that trains on its own outputs needs a rule for which outputs to trust, and aTTT's rule (novelty in the update text) is an acceptor in the series' sense.

Test-time training on released checkpoints

The other 2026 thread makes test-time training work on models that were not trained for it. ByteDance Seed's In-Place TTT (April 2026, ICLR oral) "treats the final projection matrix of the ubiquitous MLP blocks as its adaptable fast weights," so any transformer already has the parameters to update 20. TTT-NTP (June 2026, revised August 2026) supervises those in-place writes with the model's own next contextual hidden state, tying each update to next-token prediction; on RULER averaged over 4K to 32K contexts it was the only method to improve all four released backbones it tested, by 3.9 points on Llama-3.1-8B, 3.0 on Mistral-7B-v0.3, 4.1 on Qwen3-4B, and 2.9 on Qwen3-0.6B, while preserving commonsense and knowledge performance 21. TTCD (August 2026) uses a long-window teacher to supervise a short-window student's fast weights so updates retain what future predictions need; its in-place variant beat DeltaNet, Gated DeltaNet, sliding-window attention, and standard TTT when pretrained from scratch 22.

Self-Guided TTT (July 2026) addresses what to train on. On LongBench-v2, test-time training on randomly sampled spans of a long context hurt performance while training on oracle spans helped substantially; S-TTT has the model pick the evidence spans itself before adapting, and improved Qwen3-4B-Thinking and Llama-3.1-8B on LongBench-v2 and LongBench-Pro by up to 15% relative 23. The pattern repeats across all three: the gradient step is cheap, and the hard part is choosing the signal it follows.

Section 07 / 14

#Memory written into weights during the forward pass

Google's research line builds the gradient step into the architecture. Its lineage runs through Titans (December 2024), a neural long-term memory updated at every token by a gradient step on a "surprise" loss 24; ATLAS (May 2025), which optimizes that memory over windows of tokens and reports "+80% accuracy" over Titans on BABILong at 10M-token context 25; and Nested Learning's HOPE (NeurIPS 2025), a self-modifying architecture with memory blocks that update at different rates 26.

The 2026 papers in this line push toward consolidation across time. "Language Models Need Sleep" (June 2026, revised July 2026), from Ali Behrouz, Vahab Mirrokni and colleagues at Google, argues that existing models cannot "transfer their temporal in-context knowledge to their long-term parameters" and proposes a sleep phase with two stages: memory consolidation, which distills a smaller model's memories into a larger network by combining on-policy distillation with RL-based imitation, and dreaming, in which the model uses RL to generate a synthetic curriculum that rehearses new knowledge 27. Proteus (August 2026, Behrouz and Mirrokni with Mila's Aaron Courville) observes that static memory lets early tokens "pollute" the state and schedules memory capacity to grow with context length; applied to SWLA, Comba, Titans, and Hope-Attention, it improved language modeling, reasoning, and long-context retrieval, with gains that grow at longer contexts 28.

The shared limitation is scope. These models are evaluated at academic scale on language modeling and long-context retrieval, and the learning they do mostly lives within a context rather than across users and weeks. The sleep proposal is the first in the line to target Karpathy's distillation problem directly, and it reports qualitative support rather than a deployment.

Section 08 / 14

#Composing the old tools

A second research cluster keeps the standard transformer and asks whether classic anti-forgetting methods work at LLM scale when used together. Johns Hopkins' "Continual Learning Mechanisms Compose for Long-Horizon Memorization" (submitted September 7, 2026) sets a hard task: learn 100 question-answering tasks in sequence by continual SFT, with no replay buffer and no task identifiers 29. The authors organize mechanisms along two axes (anchors that decide what prior information each update preserves, whether data, function, or weights; and low-rank allocation rules that decide where updates are stored) and search combinations with task-level successive halving. The combination of all three anchors plus merged LoRA "raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement," with the data anchor and merged LoRA contributing most and interacting super-additively 29.

The same tools are showing up in production retrieval. A September 2026 study of a production skill router over 34,396 skills found that fine-tuning on synthetic data improved in-distribution retrieval but caused catastrophic forgetting on real and out-of-distribution queries 30. Regularizers borrowed from continual learning (embedding anchors, Learning without Forgetting, EWC, L2-to-initialization) kept out-of-distribution performance and also improved in-distribution retrieval by 13.98% for 0.6B Qwen retriever and reranker models 30.

Composition supports the practitioner view that forgetting is an engineering problem with partial fixes. It also shows the distance left: 34.9% retention after 100 tasks means losing about two-thirds of what was learned, on a synthetic stream with clean task boundaries.

Fig 05.1 · Anti-forgetting fixes, measuredMeta FAIR 2025 · JHU 2026
NATURALQUESTIONS F1 DROP (%) after learning new facts · lower is better 0 100 Full fine-tuning 89 LoRA 71 Sparse memory FT 11 AVERAGE FINAL RETENTION (%) 100 sequential QA tasks · higher is better 0 100 Naive sequential SFT 1.2 All three anchors + merged LoRA 34.9
Large relative gains, low absolutes. Left: after learning new facts, sparse memory finetuning loses 11% of NaturalQuestions F1, against 89% for full fine-tuning and 71% for LoRA, but it requires a memory-layer model 16. Right: on a stream of 100 QA tasks with no replay and no task IDs, composing all three anchors with merged LoRA raises average final retention from 1.2% to 34.9%, which still loses about two-thirds of what was learned 29. The two panels come from different papers and setups and use different metrics.
Section 09 / 14

#Adapters from live data

A different route writes knowledge into weights without per-task gradient descent. In September 2026, Boltzbit and José Miguel Hernández-Lobato published a working paper on an "Infinite-Parameter LLM": a compact hypernetwork turns live interaction data into low-rank modulations of a shared base network, and a Bayesian belief over the generator's latent state is updated online during a session, so the effective weights are re-derived as the belief evolves 31. The abstract argues that representing live data in weights "amortises compute, frees the context window, persists updates across turns, and can generalise better than in-context use" 31. The abstract describes an evaluation protocol for in-context learning and retrieval and reports no headline numbers, so for now it is a design rather than a result. Its predecessors, Sakana's Text-to-LoRA and Drag-and-Drop LLMs (both June 2025), generate adapters from task descriptions or prompts in one forward pass; Drag-and-Drop's claim of up to 30% gains over the strongest training LoRAs on unseen benchmarks remains unreplicated 32,33.

Mind Lab's Macaron-V1 (August 2026) applies the swappable idea at production scale: a frozen 744B GLM-5.2 base with four specialist LoRAs (chat, agent, coding, generative UI) routed per turn, with continual learning framed as adding and retraining specialists without touching the base 34. Its authors write that "compounding gains from continual learning and collective intelligence remain open questions" 34. Report 12 covers Macaron as a co-evolution system.

Section 10 / 14

#What the numbers show

Date System Setting Result Baseline / context Ref
2026-09 CL mechanisms compose 100 sequential QA tasks, no replay, no task IDs 34.9% average final retention Naive sequential SFT 1.2% 29
2026-09 Skill-router forgetting study Production router, 34,396 skills, 0.6B Qwen CL regularizers keep OOD retrieval and add 13.98% in-distribution Plain synthetic-data fine-tuning (forgets OOD) 30
2026-08 Shopify Sidekick loop GraphQL agent, up to 2,000 req/min, daily SFT + GRPO Beats frontier baseline on judge (no number given); est. serving cost $27M → ~$1M/yr (−96%) Frontier-model product 1
2026-08 Continual Reasoning Gym Sequential RLVR, five task sequences Modest forgetting; only on-policy prompt replay reaches multitask level Joint multitask RLVR 12
2026-07 When Does CL Require Learning Qwen3-8B, 8 methods, 4 families Online RL best at fact updates, fails on noisy reward; GEPA highest drift F1 (0.24) but stable F1 falls 0.32 → 0.26 Same backbone, shared protocol 7
2026-07 Agentic TTT Live LoRA updates in agent episodes +5.0 pts ALFWorld, +4.9 pts SWE-bench Lite at 1.9× cost No test-time training 19
2026-07 Self-Guided TTT LongBench-v2, LongBench-Pro Up to 15% relative Base model; random-span TTT hurts 23
2026-06 TTT-NTP RULER 4K–32K, four released backbones +2.9 to +4.1 pts on every backbone Released checkpoint 21
2026-05 Learning, Fast and Slow Reasoning tasks, shifting domains Up to 3× sample efficiency; up to 70% less KL Parameter-only RL 15
2026-03 Cursor Composer real-time RL Coding agent, as often as every five hours Edit persistence +2.28%; dissatisfied follow-ups −3.13% A/B vs prior checkpoint 2
2025-10 Sparse memory finetuning New facts, memory-layer model NaturalQuestions F1 drop 11% Full FT 89%, LoRA 71% 16
2025-09 Cursor Tab online RL 400M+ requests/day, 1.5–2 h cycle 21% fewer suggestions, 28% higher accept rate Previous Tab model 8

The production numbers are company-reported, on metrics the company chose, without confidence intervals; Shopify's quality claim has no number at all. The research numbers come from settings built to isolate one type of change. No study yet measures a weight-updating agent on a realistic deployment stream against a matched-budget text-memory baseline, which is the comparison that would settle where continual learning should live.

Section 11 / 14

#Where it breaks

The evaluator is the attack surface. Cursor's production exploits and Composer 2.5's synthetic ones came from rewards that measured a proxy for success. Shopify's loop replaces user signals with a calibrated judge, which moves the problem rather than removing it: a daily GRPO loop against a fixed judge will find that judge's blind spots, and the degradation tests Shopify runs check only the failures someone thought to simulate. Report 3 (Self-generated rewards) documents how LLM judges get gamed. Cursor caught its exploits by watching aggregate metrics; neither company describes a gate that would catch an exploit that looks like good behavior.

Self-training loops amplify drift. aTTT's finding that continuous test-time updates "amplify drift when the agent gets stuck" 19 is the per-episode version of a risk every loop that trains on its own outputs carries. Shopify's use of external critics and human annotators for the cases critics cannot repair is its defense; the Berkeley study's finding that GRPO "may reinforce the wrong answer" under noisy reward 7 is what happens without one.

What goes into weights cannot easily come out. K-Bench (September 2026) tested unlearning for agents and found that when a secret lives in the weights, "none of the twenty evaluated published methods demonstrably removes it"; when the secret sits in the prompt or retrieval store, standard benchmarks report no leakage while the deployed agent still leaks it on 22–86% of queries 35. A loop that folds production traffic into shared weights daily inherits that problem for every user request to delete data.

Moving learning to memory relocates forgetting rather than removing it. Hu, Long, and Wang's April 2026 study found that finer-grained external memory can improve forward transfer while inducing severe forgetting; the stability-plasticity dilemma moves to memory representation and retrieval 36. CL-Bench (June 2026) finds that "naive ICL outperforms systems dedicated to memory management" across six expert-validated domains 37, and AgentCL (June 2026) reports plasticity gains of 13.5 to 21.9 points on compositional code streams (up to 26.0 on BrowseComp+), stability gains between −3.5 and +4.2, and negative generalization to held-out tasks (−4.2 to −0.8) 38. Neither benchmark evaluates weight-updating agents directly (see 14, Measuring self-improvement).

Updates do not survive a base-model upgrade. A June 2026 survey of industrial continual learning names "capability inheritance breakage" on foundation-model upgrades as a core challenge, alongside plasticity erosion from repeated adaptation 39. A skill file carries over to the next model; a gradient update to last quarter's checkpoint does not. Shopify's design partly answers this: the durable asset is the judge, the repaired-trajectory corpus, and the loop, all of which can retrain a new base.

The pre-deployment checkpoint disappears. Dwarkesh Patel's August 2026 governance argument: "A lot of the proposals that have been put forward for regulating AI assume that you train a model, and then you deploy it. … But what if the base model is getting updated every single day based on the millions of sessions of work it does?" 5. CursorBench and Shopify's judge are small-scale answers: evaluate every checkpoint before it ships. That works for a coding model or a store-query agent. It is unclear what the equivalent gate is for a general model whose dangerous capabilities take days to evaluate. Patel draws a deregulatory conclusion; the same premise supports the opposite one, that evaluation has to move into the loop. Report 15 takes up oversight.

Section 12 / 14

#State of play, September 2026

The summer of 2026 changed the production picture more than the research one. Until August, Cursor was the only company with a documented weight loop from deployment; Shopify's post 1 made it two, and showed the loop can run outside coding, on a judge rather than raw user signals, as a distillation from a frontier model into a cheaper specialist. The labs remain silent.

The research record from July to September is consistent. The Berkeley comparison 7 showed that the choice between text and weights depends on how the environment changes, with RL strongest for fact updates and weakest under noise. The Continual Reasoning Gym 12 confirmed that sequential RL forgets modestly and that on-policy replay closes the rest of the gap. Test-time training moved into agent episodes 19 and onto released checkpoints 23,21,22. Google's memory-architecture line turned toward consolidation and capacity scheduling 28,27. Composition put the first large number on sequential fine-tuning 29, and K-Bench showed that deletion from weights is unsolved 35. A small infrastructure layer exists: Resolute Labs' CLaaS (June 2026) wraps asynchronous gradient updates from a replay buffer behind a chat API and, on one adversarial task, found parametric updates with replay beat in-context learning on forward transfer and forgetting 40.

The two tracks have started to meet. Both production loops are RL with replay or monitoring and a hint-conditioned teacher that turns text feedback into dense signal, which is the explore-in-context, consolidate-into-weights design the series keeps returning to. The research track supplies the reasons (on-policy updates forget less, replay closes the gap, noisy rewards break RL) but not yet the evaluations at deployment scale.

Section 13 / 14

#What ships

Weight-level continual learning ships with documentation in two places: Cursor's Tab and Composer models, with new checkpoints every 1.5–2 hours for Tab and as often as every five hours for Composer 8,2, and Shopify's Sidekick GraphQL agent, retrained daily 1. Macaron-V1 ships modular LoRA specialists on a frozen base without describing updates from user traffic 34. Boltzbit's live-data adapter is a working paper 31. Open-source code exists for In-Place TTT and TTT-NTP 20,21, but both remain research-grade. Several startups market "continual learning"; Synth AI's continual-learn.com turns out to optimize prompts and context with MIPRO rather than update weights 41, so the label alone says little about where learning is stored. Most teams that want agents to improve from deployment still do it in text, and report 16 (What ships) covers that landscape.

Section 14 / 14

#Open problems

The lab-product gap. Both documented loops run on narrow products with dense, domain-specific signals. Whether frontier chat and agent models can get a comparable signal, and whether anyone will accept a shared model that changes daily, decides whether continual learning becomes the default or stays a technique for specialists distilled from frontier models.

Where the evaluator lives when the model never stops training. Cursor's exploits were caught by humans watching aggregate metrics; Shopify's judge is checked against past A/B tests and targeted degradations. A continual loop needs an acceptor that runs as often as the update and catches reward hacks that look like good behavior. Until that exists, every production loop depends on how closely its operators watch.

Long-horizon retention and deletion. 34.9% retention after 100 clean tasks still means forgetting most of what was learned, and no published unlearning method demonstrably removes a secret from weights. A loop that learns from users daily needs both to keep what matters and to remove what must go.

Scaling test-time memory. The memory architectures and test-time training methods are validated at 8B parameters or fewer. Nobody has shown that per-chunk weight updates keep their advantage, or stay affordable at inference, at frontier scale; the in-place methods are the most promising path because they need no new pretraining.

Three claims survive the evidence. Weight-level learning after deployment works in at least two products and pays for itself, in Shopify's case mainly through cost. It is an RL problem more than a fine-tuning problem, because on-policy updates with replay forget least, and it becomes tractable when a text hint turns sparse feedback into a demonstration. Its binding constraint is the same one that bounds every loop in this series: the evaluator, whether users or a calibrated judge, which the model will search as hard as it searches the task.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]A. McNamara and C. Mazza-Anthony, “Sidekick's continual learning loop,” Shopify Engineering, 2026-08-05. https://shopify.engineering/sidekicks-continual-learning-loop
  2. [2]Cursor (J. Jackson, B. Trapani, N. Wang, W. Zhu), “Improving Composer through real-time RL,” Cursor blog, 2026-03-26. https://cursor.com/blog/real-time-rl-for-composer
  3. [3]A. Karpathy and D. Patel, “Andrej Karpathy — AGI is still a decade away,” Dwarkesh Podcast, 2025-10-17. https://www.dwarkesh.com/p/andrej-karpathy
  4. [4]D. Silver and R. S. Sutton, “Welcome to the Era of Experience,” Google DeepMind position paper, April 2025. https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf
  5. [5]D. Patel, “8 Predictions for the Era of Continual Learning,” dwarkesh.com, 2026-08-07. https://www.dwarkesh.com/p/era-of-continual-learning
  6. [6]D. Amodei and D. Patel, Dwarkesh Podcast interview, February 2026. https://www.youtube.com/watch?v=mYDSSRS-B5U
  7. [7]A. Harrington, N. Saxena, M. Murphy, A. Borovykh, Z. Yun, S. Kamath, A. E. Kyi, T. Darrell, J. Malik, Y. Bai, “When Does Continual Learning Require Learning,” arXiv:2607.07847, 2026-07-08. https://arxiv.org/abs/2607.07847
  8. [8]Cursor, “Improving Cursor Tab with online RL,” Cursor blog, 2025-09-12. https://cursor.com/blog/tab-rl
  9. [9]Cursor, “Composer 2 Technical Report,” arXiv:2603.24477, March 2026. https://arxiv.org/abs/2603.24477
  10. [10]Cursor, “Composer 2.5,” Cursor blog, 2026-05-18. https://cursor.com/blog/composer-2-5
  11. [11]C. Gou, H. Tu, Y. Fang, J. Cai, H. Rezatofighi, “Sample-Efficient Learning from Agent Experience,” arXiv:2607.21051, 2026-07-23. https://arxiv.org/abs/2607.21051
  12. [12]L. Luo, G. Zhang, H. Xu, R. Li, C. Fang, L. Fan, “Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR,” arXiv:2608.18574, 2026-08-19. https://arxiv.org/abs/2608.18574
  13. [13]I. Shenfeld, J. Pari, P. Agrawal, “RL's Razor: Why Online Reinforcement Learning Forgets Less,” arXiv:2509.04259, 2025-09-04 (ICLR 2026). https://arxiv.org/abs/2509.04259
  14. [14]I. Shenfeld, M. Damani, J. Hübotter, P. Agrawal, “Self-Distillation Enables Continual Learning,” arXiv:2601.19897, 2026-01-27. https://arxiv.org/abs/2601.19897
  15. [15]R. Tiwari et al., “Learning, Fast and Slow: Towards LLMs That Adapt Continually,” arXiv:2605.12484, 2026-05-12. https://arxiv.org/abs/2605.12484
  16. [16]J. Lin et al., “Continual Learning via Sparse Memory Finetuning,” FAIR at Meta, arXiv:2510.15103, 2025-10-16. https://arxiv.org/abs/2510.15103
  17. [17]E. Akyürek et al., “The Surprising Effectiveness of Test-Time Training for Few-Shot Learning,” arXiv:2411.07279, 2024-11-11 (ICML 2025). https://arxiv.org/abs/2411.07279
  18. [18]Y. Zuo, K. Zhang et al., “TTRL: Test-Time Reinforcement Learning,” arXiv:2504.16084, 2025-04-22. https://arxiv.org/abs/2504.16084
  19. [19]Y. Wang, J. Hao, Y. Shi, K. Yuan, M. Sun, “No Time Like the Present: Agentic Test-Time Training for LLM Agents,” arXiv:2607.03441, 2026-07-03. https://arxiv.org/abs/2607.03441
  20. [20]G. Feng et al., “In-Place Test-Time Training,” arXiv:2604.06169, 2026-04-07 (ICLR 2026 oral). https://arxiv.org/abs/2604.06169 · code: https://github.com/ByteDance-Seed/In-Place-TTT
  21. [21]X. Ouyang, Z. Cai, J. Hu, “Test-Time Training with Next-Token Prediction,” arXiv:2606.21803, 2026-06-19 (v2 2026-08-30). https://arxiv.org/abs/2606.21803
  22. [22]Z. Wang et al., “Learning What to Remember: Test-Time Training via Context Distillation,” arXiv:2608.01672, 2026-08-03. https://arxiv.org/abs/2608.01672
  23. [23]X. Zhu, Z. Xu, X. Wei et al., “Self-Guided Test-Time Training for Long-Context LLMs,” arXiv:2607.09415, 2026-07-10. https://arxiv.org/abs/2607.09415
  24. [24]A. Behrouz et al., “Titans: Learning to Memorize at Test Time,” arXiv:2501.00663, 2024-12-31. https://arxiv.org/abs/2501.00663
  25. [25]A. Behrouz et al., “ATLAS: Learning to Optimally Memorize the Context at Test Time,” arXiv:2505.23735, 2025-05-29. https://arxiv.org/abs/2505.23735
  26. [26]A. Behrouz, M. Razaviyayn, P. Zhong, V. Mirrokni, “Nested Learning: The Illusion of Deep Learning Architectures,” NeurIPS 2025; arXiv:2512.24695. https://arxiv.org/abs/2512.24695
  27. [27]A. Behrouz, F. Hashemi, A. Javanmard, V. Mirrokni, “Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories,” arXiv:2606.03979, 2026-06-02 (v2 2026-07-10). https://arxiv.org/abs/2606.03979
  28. [28]R. Bayat, A. Behrouz, V. Mirrokni, A. Courville, “Proteus: Incremental Memory Activation for Long-Context Sequence Modeling,” arXiv:2608.16844, 2026-08-17. https://arxiv.org/abs/2608.16844
  29. [29]Z. Zhang, A. Zhang, D. Khashabi, T. Shu, “Continual Learning Mechanisms Compose for Long-Horizon Memorization,” arXiv:2609.06986, 2026-09-07. https://arxiv.org/abs/2609.06986
  30. [30]S. S. Murtaza, Y. Nie, U. Soni, E. Wen, A. Frydenlund, “When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents,” arXiv:2609.10750, 2026-09-09. https://arxiv.org/abs/2609.10750
  31. [31]J. Hu, R. M. Clarke, Y. Zhang, J. M. Hernández-Lobato (Boltzbit), “Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data,” arXiv:2609.18842, 2026-09-16 (v2 2026-09-21). https://arxiv.org/abs/2609.18842
  32. [32]Sakana AI, “Text-to-LoRA: Instant Transformer Adaption,” arXiv:2506.06105, June 2025. https://arxiv.org/abs/2506.06105
  33. [33]Z. Liang et al., “Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights,” arXiv:2506.16406, 2025-06-19. https://arxiv.org/abs/2506.16406
  34. [34]Mind Lab, “Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA,” arXiv:2608.09819, 2026-08-10. https://arxiv.org/abs/2608.09819
  35. [35]G. Yu, Y. Jiang, Q. Wang, B. Ma, X. Wang, “K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments,” arXiv:2609.12808, 2026-09-11. https://arxiv.org/abs/2609.12808
  36. [36]Q. Hu, Q. Long, W. Wang, “When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents,” arXiv:2604.27003, 2026-04-29. https://arxiv.org/abs/2604.27003
  37. [37]“Continual Learning Bench (CL-Bench): Evaluating Frontier AI Systems in Real-World Stateful Environments,” arXiv:2606.05661, June 2026. https://arxiv.org/abs/2606.05661
  38. [38]Y. Shu et al., “AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents,” arXiv:2606.02461, 2026-06-01. https://arxiv.org/abs/2606.02461
  39. [39]S. Li et al., “LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning,” arXiv:2606.24901, June 2026. https://arxiv.org/abs/2606.24901
  40. [40]K. Fallah, S. Naihin, B. Widawsky, Q. Mao, “CLaaS: Continual learning as a service for sample efficient online learning,” arXiv:2606.05559, 2026-06-04. https://arxiv.org/abs/2606.05559
  41. [41]Synth AI, continual-learn.com product site, accessed September 2026. https://www.continual-learn.com/