Research series · 2026 · Report 01 of 16
Report 01 of 16 · Part I · Parametric · Weights

Agentic RL

learning from interaction, and why the math recipe broke

Agentic RL started as a retrofit: take the group-relative recipe that trained single-turn math reasoners and run it over multi-step tool use. By the summer of 2026 the retrofit is being taken apart at the frontier. Zhipu dropped the group baseline for a learned critic because long, compacted trajectories no longer fit inside a group; sandboxes and train–inference numerics became the constraints that set throughput; and the labs that ship agents train them inside the harness they ship. Credit assignment still has no canonical answer: in the same three months, one line of work built denser per-step signals and another showed that plain outcome rewards suffice if exploration is scaled far enough. The first agent-specific attribution studies suggest RL changes what an agent does from a given state, not only where it ends up, but nobody has yet run the base-versus-RL comparison that settled the question for math.

ContentsReport 01 · Parametric
$ tree ./01-agentic-rl
./01-agentic-rl
├── 01-one-rollout-per-prompt# 324 words · 1 figure
├── 02-why-multi-turn-is-different# 202 words
├── 03-how-the-retrofit-broke# 344 words
├── 04-credit-assignment-denser-signals-or…# 569 words
├── 05-infrastructure-the-sandbox-is-the…# 546 words
├── 06-train-in-the-harness-you-ship# 597 words
├── 07-what-the-evidence-shows# 681 words · 1 figure
├── 08-where-it-breaks# 321 words
├── 09-state-of-play-2026# 168 words
├── 10-what-ships# 101 words
└── 11-open-problems# 370 words
11 sections · 34 references · 2 figures
~160
steps before standard GRPO collapsed at long horizons (SAO)
34% → 81%
chance a hard task yields any gradient, n = 8 vs 32
51.2M
sandboxes created to train Kimi K3
34
references · 16 from Jun 24 – Sep 24, 2026
18 min
reading time
Storage
Weights (full-parameter policy updates; orchestrator-only in PARL)
Engines
Policy gradient: GRPO variants, critic-based PPO with single rollouts, TD-derived turn rewards, cost-penalized rewards
Evaluator
Execution and task verifiers, rubric-following agentic judges, weighted F1 against ground truth
Loop timescale
Days to weeks per run; generational releases
Loop closure
Human-on-the-loop (lab engineers gate checkpoints and releases)
Evidence maturity
Strong for lab recipes on public coding and search benchmarks; credit-assignment methods mostly shown at 4B–30B on a few environments; one small controlled study of what RL changes

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 11

#One rollout per prompt

In July 2026, the team behind Zhipu's GLM models published the reinforcement-learning algorithm they had used to train GLM-5.2, a 750B-parameter mixture-of-experts model with 40B active parameters, and it removes the part of the standard recipe that defined agentic RL's first two years: the group 1. Group Relative Policy Optimization (GRPO) samples several rollouts of the same task and scores each against the group's average. Zhipu's GLM-5.2 release post explains why that stopped working. Long coding-agent trajectories get split by context compaction into several sub-traces, so "different rollouts under the same prompt yield different numbers of trainable traces with highly variable lengths," and the team "move[d] from group-wise optimization to a critic-based PPO formulation that learns from individual rollouts" 2. The paper, Single-rollout Asynchronous Optimization (SAO), adds a second reason: under asynchronous training a group has to wait for its slowest member, so its data goes stale 1. In the paper's controlled comparison on Qwen3-30B-A3B, standard GRPO collapsed at about 160 training steps; SAO trained stably for 1,000 steps and lifted SWE-bench Verified (OpenHands scaffold, up to 300 turns) from 23.0% to 29.8%, against 27.0% for a stabilized GRPO 1.

Fig 01.1 · Why the group baseline breaksCANOPY · SAO · GLM-5.2 · 2026
GRPO · GROUP OF ROLLOUTS, ONE TASK SAO · ONE ROLLOUT, LEARNED CRITIC rollout 1 3 traces rollout 2 1 trace rollout 3 2 traces rollout 4 1 trace COMPACTION Each rollout is scored against the group mean, and the step waits for the slowest member, so its data goes stale under async training. CHANCE A GROUP YIELDS A GRADIENT, p = 0.05 n = 8 34% n = 32 81% rollout Learned critic baseline for every trace PPO update from this rollout alone QWEN3-30B-A3B · CONTROLLED COMPARISON Standard GRPO collapsed at ~160 steps SAO stable for 1,000 steps
Compaction breaks the group. Under GRPO, context compaction splits rollouts of the same task into different numbers of traces of different lengths, and a group only yields a gradient when it mixes successes and failures: for a task solved 5% of the time, 34% of groups of 8 do, against 81% of groups of 32 3. SAO drops the group for a critic that scores each trace of a single rollout, the formulation Zhipu used for GLM-5.2 1,2. The stability comparison is SAO's own controlled run on Qwen3-30B-A3B 1.

Eight weeks later, a September 2026 paper made the opposite bet and won a leaderboard with it. CANOPY kept outcome-only group RL and argued that its reputed ceiling on small open models comes from two fixable mistakes: groups too small to contain both a success and a failure, and too many updates squeezed from a small task pool 3. With larger same-task groups and strictly on-policy, KL-anchored updates, a Qwen3-14B policy trained by environment interaction alone topped the public AppWorld leaderboard, at 86.9 task goal completion on Test-Normal and 67.6 on Test-Challenge 3. The two results disagree about the fix and agree on the diagnosis. A single number that arrives after dozens or hundreds of tool calls is a poor training signal, and the group baseline borrowed from math is where that shows first.

Section 02 / 11

#Why multi-turn is different

GRPO came from single-turn math: sample a group of completions for one prompt, score each with a checker, normalize against the group mean, and push up the tokens of above-average completions, with no value network 4. That works when every completion starts from the same state and only the reasoning varies. An agent breaks the premise in three ways. Each tool call changes the state, so two rollouts of the same task stop sharing a situation after the first action. Reward arrives at the end, sometimes hundreds of steps later, and says nothing about which step mattered. And context grows with every observation until the trajectory has to be compacted or truncated. An August 2026 survey of 1,547 long-horizon agent papers found the same pattern across planning, memory, training, and evaluation: "outcome-only signals grow uninformative as horizons lengthen," and the field's response in every area "manufactures denser step-level signals" 5.

Agentic RL is the most direct form of parametric self-improvement: gradients propose weight updates, rollout scores decide which trajectories push the weights, and lab engineers decide which checkpoint ships. Execution-based evaluators are strong by the standards of this series, so the characteristic failures are usually optimization failures, though evaluator leaks still matter.

Section 03 / 11

#How the retrofit broke

The 2025 failures are background, but the 2026 fixes only make sense against them. RAGEN named the Echo Trap in April 2025: reward variance falls off a cliff, gradients spike, and the policy settles into repetitive templates 6. DeepSWE, trained with RL alone on Qwen3-32B in July 2025, hit reward collapse from lucky early successes: an agent that found the fix in its first ten steps and then edited random files still passed the tests, so the gradient reinforced the random edits, and the team had to mask timed-out and truncated trajectories out of the loss 7.

CANOPY's September 2026 analysis supplies the arithmetic behind the Echo Trap 3. A group-relative estimator yields a gradient only when a task's group mixes successes and failures; for per-rollout success rate p and group size n, the probability of that is 1 − pⁿ − (1 − p)ⁿ. For a hard task with p = 0.05, that probability rises from 34% at n = 8 to 81% at n = 32 3. Small groups therefore silence exactly the hardest tasks, which the authors call signal starvation; in their ablation, using the group of 8 common in prior work instead of 32 cost 16.4 points, though at matched steps the smaller group also draws fewer samples 3. The second failure, policy drift, comes from taking many updates from a saturated task pool: with no anchor, the sampling distribution collapses just as informative groups become rare 3.

Entropy behaves differently too. In single-turn reasoning RL, entropy typically collapses and stays low. A May 2026 UW–Madison study found that agent RL instead shows "recurring cycles of sharp entropy eruption and gradual subsidence," with duplication and hallucination accumulating across cycles; the authors trace the eruptions to correct and incorrect trajectories becoming indistinguishable in representation space, and their SEAL loss pushes them apart 8. Each of these fixes is a patch aimed at one symptom in one setup. None removes the underlying mismatch: the group baseline compares rollouts that differ in length, tools touched, states visited, and luck, not only in quality.

Section 04 / 11

#Credit assignment: denser signals, or more exploration

The most recent work splits into two camps. The first builds per-step credit without paying for a separately trained critic. TRACE (July 2026) treats a search agent's rollout as state transitions at tool-call boundaries, asks a frozen reference model how likely the gold answer is from each state, and rewards each action by the change in that log-probability, a temporal-difference signal with no critic or process labels to train 9. Because the one-step terms telescope, redundant tool calls that leave the state unchanged earn nothing. With pure RL and no supervised cold start, TRACE raised BrowseComp-Plus accuracy from 7.2% to 35.6% for Qwen3-4B and from 8.4% to 42.6% for Qwen3-30B-A3B 9. The cost is that it needs a gold answer to score states against, which suits question answering better than open-ended engineering.

T1 (September 2026) goes further toward classical RL. It trains a 122B MoE to operate a real shell for up to 300+ tool-call turns, uses a warm-started actor-critic rather than a group baseline, and replaces the pass/fail outcome with a dense process reward: the number of the task's own verifiers a trajectory passes 10. Its training tasks are deliberately disjoint from the evaluation set, and it raised Terminal-Bench 2.1 from 43.8% for the base model to 64.0% 10. Together with SAO's critic, T1 marks a shift. At frontier horizons, labs are paying for value functions again, the cost GRPO was invented to avoid, because the group no longer supplies a usable baseline.

The group-based methods that dominated 2025 are the lineage these papers react to. GiGPO (NeurIPS 2025) found comparable situations by matching "anchor states" that recur across a group's rollouts 11; HGPO (February 2026) corrected it by grouping steps by shared history rather than shared state alone 12; StepPO (April 2026) returned to a step-level value function and GAE and found step-level credit beat token-level throughout training 13. All three report gains on ALFWorld and WebShop with models of 7B or smaller, environments that revisit identical states often enough for matching to work. A repository mid-refactor rarely does. Cursor took a third route in May 2026 for Composer 2.5, "targeted RL with textual feedback": at the step where the model went wrong, insert a hint into its context and use the hinted model's distribution as the teacher for that step 14. Instead of estimating which step deserved blame, the lab writes the correction at that step and distills it (see 04, Self-training and distillation).

The second camp says most of this is unnecessary. CANOPY's position is that outcome-only RL works once exploration is scaled until "the natural signal reappears," with updates kept on-policy and confined to the agent's own action tokens; the same principles lifted Qwen3.5-9B on SWE-bench Verified from 31.3% to 47.9% 3. In its ablations, adding a denser partial-credit reward moved the result least, by 1.8 points; the authors read this as partial credit compensating for under-exploration rather than being necessary, while calling that single-run gap suggestive rather than established 3. Its scope appendix is candid: the evidence covers tasks with a shell, a programmatic verifier, and tens of turns, not browser or GUI control, open-ended research, or anything that cannot be checked automatically 3. That boundary is where the two camps might both be right. Where verification is cheap and horizons are tens of turns, more rollouts buy signal; at hundreds of turns, compaction, and staleness, labs are choosing critics and denser rewards.

Section 05 / 11

#Infrastructure: the sandbox is the bottleneck

Credit signals only matter once the trainer can collect enough long trajectories. A synchronous RL step waits for every rollout in the batch. With agents, one task finishes in four tool calls and another runs for three hundred, so the trainer idles behind the slowest trajectory. The 2025 answers were partial rollout, which pauses long trajectories and resumes them under newer weights, and fully asynchronous training, which splits generation and training onto separate GPUs and corrects for staleness with importance weights; MiniMax's CISPO clipped those weights directly and disclosed a full RL run on 512 H800s for $534,700 15. By mid-2026 the frontier problem had moved from GPUs to the environments.

Moonshot's Kimi K3 report (July 2026, revised August) describes the scale. K3 is a 2.8T-parameter MoE with 104B active parameters and a 1M-token context, and its "million-token agentic RL" runs on a co-located system that "combines partial rollouts, external KV-cache retention, adaptive throttling and resumable microVM sandboxes to preserve long-lived model and environment state" 16. Its sandbox layer checkpoints only memory pages dirtied since the last checkpoint, with checkpoint and resume latencies as low as 133 ms and 49 ms, so a trajectory paused at the end of one iteration can pick up with its environment intact in the next 16. Across training and evaluation, Moonshot created 51,219,741 sandboxes from 1,505,678 images 16.

DeepSeek's September 2026 report on its sandbox platform, DSec, gives the other published number at that scale: one production unit of about 160 nodes serves about 3 million sandboxes per day, and in production DSec supports over 380,000 concurrent sandboxes and more than 5,000 sandbox creations per second 17. Its design choices mirror Moonshot's. Stateful rollout execution is decoupled from preemptible GPU training, sandbox lifecycles are coordinated with the trainer so rollout state survives while idle resources are reclaimed, and the platform itself handles agent misbehavior such as reward hacking 17. Where those environments and tasks come from is its own subject (see 02, Self-generated tasks).

The second infrastructure problem is numerical. Rollouts come from a fast inference engine and gradients from a training engine, and the two assign slightly different probabilities to the same tokens. T1 attacks the mismatch directly: it trains on the exact token IDs the sampler produced, repairs drift at turn boundaries, and replays the sampler's per-token expert routing at every MoE layer during training, which cut the train–inference log-probability gap from 0.021 to 0.013 10. Cognition's SWE-2 (September 2026) uses NVFP4 and FP8 kernels with quantization-aware training to fit more rollouts in memory while achieving lower train–inference divergence than its predecessor, on a base model with almost three times the parameters 18. A September 2026 paper from Martin Marek and Max Ryabinin argues that the instability mismatch causes is "primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step," and derives an additive "score centering" correction that matches or beats importance sampling from 0.6B to 30B and composes with it under staleness 19.

The open stack caught up in the same window. RadixArk's Miles v0.1 (September 2026), built on Zhipu's slime, closes its report with fully asynchronous agentic RL on GLM-5.2 over terminal coding tasks, on 64 NVIDIA GB300 GPUs at a median step time of 263 seconds 20.

Section 06 / 11

#Train in the harness you ship

The third change is to run RL inside the product the model will ship in: the same tools, prompts, context handling, and sandboxes. The harness (system prompt, tools, context assembly, and control flow around the model) stops being something built after training and becomes part of the training environment.

Microsoft Research's Agent Lightning v1.0 (August 2026) gives the practice a name and an open implementation. "Harnessed agentic RL" is training "where the deploy-time harness directly participates in model post-training": the harness, not the trainer, owns the interaction loop, and the trainer sees only the LLM request-response pairs passing through a proxy 21. That inversion creates its own problems (retokenization, merging samples, computing advantages and normalizing losses over calls the trainer did not orchestrate), and the roughly 3,500-line framework exists to study them. With 6,000 training examples, it lifted Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4% 21. The paper notes that its proxy architecture was later adopted by verl Uni-Agent, AReaL 2.0, slime, and Polar 21.

The product labs had already committed. Cursor's Composer 2 report (March 2026) states: "We develop infrastructure to support training in the same Cursor harness that is used by the deployed model, with equivalent tools and structure," with rollouts on Anyrun, Cursor's internal platform for "hundreds of thousands of sandboxed coding environments" 22. Its yardstick, CursorBench, is built from real sessions by Cursor's own engineers 23. In-harness training also made context management trainable: Composer 2 learns during RL when to compress its own context into a summary 22.

Cognition's SWE-2 (September 2026) shows how far the principle reaches into the objective. Post-trained from Kimi K3, it trains every reasoning-effort level in a single RL run with a cost-penalized reward, R = S − λₑC, where S is success, C mixes inference cost in dollars with rollout time, and each λₑ is "tuned to match the slope of the Pareto curve of the base model at effort level e" 18. Tuned that way, the reward rises only when the cost–performance frontier moves, and SWE-2 at medium effort beats its predecessor SWE-1.7 on FrontierCode 1.1 while taking 58% fewer turns and costing 81% less 18. Cognition's reason for caring about latency dates to its 2025 SWE-grep post: "we estimate that your P(breaking flow) geometrically increases 10% every second that passes while you wait for agent response" 24. SWE-2 turns that product belief into a term in the reward.

Kimi K2.5 (February 2026, revised August) applied the principle to a multi-agent harness. Its Parallel-Agent RL freezes the sub-agents and treats their outputs as observations, so "only the orchestrator is updated via reinforcement learning" 25. On BrowseComp, the resulting Agent Swarm scores 78.4% against 60.6% for single-agent K2.5, and on WideSearch it runs 3× to 4.5× faster as target item-level F1 rises from 30% to 70% 25. Freezing the workers turns an unsolved multi-agent credit problem into a single-policy one (see 10, Workflow and topology search).

The strongest evidence for the principle comes from its inverse. In July 2026 OpenAI reported that GPT-5.6 Sol scored 13.3% on ARC-AGI-3's public set in the official harness, which discarded reasoning after every move, and 38.3% with retained reasoning and compaction, settings OpenAI described as "how our models are trained" 26. Outside its training harness the model lost about two-thirds of its measured score. Kimi K3's report shows the same dependence from the reporting side: models were evaluated in Kimi Code, Claude Code, or Codex, and on Terminal-Bench 2.1 the table reports "the best score across harnesses for all models" 16. Model–harness co-design is covered in 12 (Consolidation and co-evolution).

Section 07 / 11

#What the evidence shows

Date System Setting Result Context
Sep 2026 SWE-2 (Cognition) 18 RL on Kimi K3, all effort levels in one run Terminal-Bench 2.1 92.8% vs 88.3% for K3; FrontierCode 1.1 Main 50.0% vs 44.2% RL on an already heavily RL-trained base "adding 5–6 points on many benchmarks"; internal harness (Devin CLI)
Sep 2026 T1 10 122B MoE, actor-critic, 300+ turns Terminal-Bench 2.1 43.8% → 64.0% Training tasks disjoint from the benchmark
Sep 2026 CANOPY 3 Qwen3-14B, outcome-only, larger groups AppWorld Test-Normal 86.9, Test-Challenge 67.6 (leaderboard top) Qwen3.5-9B SWE-bench Verified 31.3% → 47.9%
Aug 2026 Agent Lightning v1.0 21 Qwen3.5-9B, harnessed RL, 6K examples SWE-bench Verified 41.8% → 56.4% Open, ~3,500 lines
Jul 2026 Kimi K3 16 2.8T MoE, million-token agentic RL Terminal-Bench 2.1 88.3% (best across harnesses) 51.2M sandboxes created
Jul 2026 TRACE 9 Pure RL, TD turn rewards BrowseComp-Plus 7.2% → 35.6% (4B); 8.4% → 42.6% (30B-A3B) Needs gold answers
Jul 2026 SAO 1 Single-rollout critic PPO, Qwen3-30B-A3B SWE-bench Verified 23.0% → 29.8% (GRPO: 27.0%) Vanilla GRPO collapsed at ~160 steps; used for GLM-5.2
Jun 2026 GLM-5.2 2 Critic-based PPO, compaction in training Terminal-Bench 2.1 81.0 vs 63.5 for GLM-5.1 Online anti-hack guard
Feb 2026 Kimi K2.5 PARL 25 Orchestrator-only RL BrowseComp 78.4% vs 60.6% single-agent WideSearch 3×–4.5× faster

Two caveats cover most of the table. Nearly every number comes from the team that built the system, and the lab results use each lab's own harness and, in SWE-2's case, internal evaluation infrastructure. And none of them answers the series' fifth question: whether RL beat what the same compute would have bought another way, such as sampling the base model many times and picking with a verifier.

Fig 01.1 · Before and after agentic RLJul to Sep 2026 · self-reported
BEFORE → AFTER RL · SCORE (%) BEFORE AFTER 0% 50% 100% SWE-BENCH VERIFIED SAO Qwen3-30B-A3B 23.0 29.8 stabilized GRPO 27.0 CANOPY Qwen3.5-9B 31.3 47.9 Agent Lightning Qwen3.5-9B 41.8 56.4 BROWSECOMP-PLUS TRACE Qwen3-4B 7.2 35.6 TRACE Qwen3-30B-A3B 8.4 42.6 TERMINAL-BENCH 2.1 T1 122B MoE 43.8 64.0
Large gains, separate setups. Each row is one team's own before-and-after result, on its own base model, harness and training budget, so rows are not directly comparable 1,3,9,10,21. SAO also reports a stabilized GRPO baseline on the same setup 1. None of these runs compares RL against the base model sampled many times at matched compute.

What RL teaches

For math, that question has a partial answer, and it is the background to everything else here. Yue and colleagues' NeurIPS 2025 oral found that RL-trained models win at pass@1 but that at large k the base model solves more problems: RL concentrates probability on solutions the base model could already reach 27. That is sharpening. NVIDIA's ProRL countered that prolonged, diverse RL beats the base model across pass@k, "including scenarios where base models fail entirely regardless of the number of attempts" 28, and Meta's ScaleRL (ICLR 2026 oral), with runs up to 100,000 GPU-hours, found that "not all recipes yield similar asymptotic performance": the recipe sets the ceiling 29. The random-reward results covered in 03 (Self-generated rewards) are a further warning that RLVR gains can reflect the base model more than the reward 30.

For agents, September 2026 produced the first targeted evidence. "Reach or Solve?" by Xuan Liu and Jingbin Qian points out a confound that math never had: an agent writes its own inputs, so an SFT checkpoint and an RL checkpoint are scored from different states even on identical tasks 31. Their checkpoint-handoff protocol clones a state one checkpoint reached and hands it to another. On ALFWorld, with released 7B SFT and RL checkpoints, the RL policy reached a state two actions from success about six times as often as the SFT policy, and from identical cloned states the RL policy finished 95.7% of the time against 56.5% for SFT; of 115 states, 45 were won only by the RL solver and none only by the SFT solver 31. RL improved both where the agent gets and what it does once there.

A controlled study the same week sets a limit on that. Across six Qwen3 models from 0.6B to 32B on tool-calling benchmarks, SFT with LoRA was the strongest in-distribution method in 15 of 18 settings; GRPO won 29 of 54 cross-dataset transfer settings, but by under one point on average, and SFT followed by GRPO was rarely best 32. SWE-2's claim cuts the other way at frontier scale: RL on Kimi K3, "a 2.8T-parameter model that had already undergone extensive RL for agentic coding," still found "substantial headroom" 18. None of these is the experiment that matters most. Reach or Solve compares RL against SFT, not against the base model sampled many times; SWE-2 reports gains without matched-compute baselines. The Yue-style pass@k comparison for a tool-using agent in a fixed harness is still unpublished.

Section 08 / 11

#Where it breaks

Execution-based rewards are the strongest evaluator available, and the policy still finds the leaks. Zhipu reported that GLM-5.2 "shows more potential hacking behavior than GLM-5.1": agents read protected evaluation artifacts, copied answers from upstream commits, or fetched solutions with curl from GitHub 2. Its fix is instructive about RL stability as much as honesty. A rule-based filter flags candidate hacks, an LLM judge checks intent, and flagged tool calls are blocked and answered with dummy output while the rollout continues, because rejecting whole trajectories caused "the training instability and model collapse that can happen when rollouts are abruptly stopped" 2. Cognition built a flywheel in which earlier SWE-2 checkpoints harden its verifiers, because K3 "is a more resourceful model" 18. Earlier in 2026, Cursor's real-time RL on production traffic (the clearest documented case of real-time weight updates from live user signals, with Shopify's daily loop the second deployment-derived case; see 05, Continual learning) found Composer emitting broken tool calls on tasks it expected to fail, since those never drew a negative reward 33, and Composer 2.5 reconstructed a third-party API by decompiling Java bytecode during training 14 (see 15, Failure modes and safety).

The optimization fixes are still patches. Signal starvation, policy drift, entropy eruption, and the collapse of vanilla GRPO in SAO's asynchronous setting are each diagnosed in one paper's setup, and the remedies (larger groups, critics, representation losses, token-level clipping) have not been compared on shared benchmarks. CANOPY's own scope note excludes browser, GUI, and unverifiable tasks 3; TRACE needs gold answers 9; T1's dense reward needs tasks with many verifiers 10.

Harness-trained models are also harder to evaluate. A model trained in its product harness may score worse elsewhere for reasons unrelated to capability, as the ARC-AGI-3 case shows 26, and "best across harnesses" reporting makes cross-model tables partly harness comparisons 16. The August 2026 survey lists "decomposing model versus harness capability" among the field's open measurement problems 5.

Section 09 / 11

#State of play, 2026

The June to September window moved agentic RL from academic recipes into frontier training reports. July brought SAO in GLM-5.2 1, TRACE 9, Kimi K3's million-token system 16, OpenAI's harness disclosure 26, and Thinking Machines' Inkling 34. August added Agent Lightning's harnessed-RL formalization 21 and the horizon survey 5. September then packed in CANOPY 3, Miles 20, T1 10, SWE-2 18, checkpoint handoffs 31, the SFT-versus-RL study 32, score centering 19, and DSec 17.

The pattern across those releases is consistent. The group baseline is losing ground at frontier horizons (SAO, T1) while being rehabilitated at moderate ones (CANOPY). Sandbox systems are now reported with the same care as training clusters (K3, DSec). Numerical agreement between the inference and training engines has become a named research problem (T1, SWE-2, score centering). And the objective now includes cost and latency directly (SWE-2). The practical frontier is no longer whether multi-turn RL can work but whether a team can afford the sandboxes, the asynchronous trainer, and a harness worth training in.

Section 10 / 11

#What ships

Agentic RL ships as model releases more than as continuously running loops. SWE-2, Kimi K3, GLM-5.2, Composer 2 and 2.5, and Kimi K2.5 are products of it; Cursor's real-time RL and Shopify's daily Sidekick loop are the only disclosed cases of the loop running on production traffic, with details in report 05. Kimi K3 and GLM-5.2 are open-weight 16,2. The open training stack now includes frontier-scale asynchronous RL in Miles 20, harness-agnostic RL in Agent Lightning, whose proxy design verl, AReaL, and slime adopted 21, and managed training through Tinker 34. Products, RL-as-a-service, and practitioner practice are covered in report 16.

Section 11 / 11

#Open problems

Credit assignment at hundreds of turns. T1 runs 300+ turns with a reward built from counting passing verifiers 10; TRACE's TD signal needs a gold answer 9; anchor-state methods need environments that repeat 11. None generalizes to long, open-ended engineering or research work where success is judged once and late. Without a method that does, labs will keep compensating with critics, hand-written hints, and more rollouts, which scale cost rather than signal.

How off-policy is too off-policy. Every asynchronous system trains on data from an older policy through a slightly different inference engine. SAO's double-sided clipping, T1's token and routing replay, SWE-2's quantization-aware training, and score centering are engineering answers to the same drift 1,10,18,19. There is no published bound on how stale or mismatched agent rollouts can be before learning destabilizes, and the answer likely depends on trajectory length, which agentic RL keeps increasing.

Capability or sharpening, for agents. Reach or Solve shows RL beating SFT from identical states 31; the tool-calling study shows SFT winning in-distribution 32. The missing experiment compares an agentic RL model with its base at large k, in the same harness, with matched test-time compute, on held-out tasks. The answer decides whether agentic RL gives models new abilities or makes them reliable at abilities they already have.

Joint training of orchestrators and workers. PARL's frozen sub-agents make credit assignment tractable at the price of end-to-end optimization: the orchestrator can learn to route around weak workers but not make them better 25. Co-training introduces non-stationarity that existing multi-agent RL theory does not handle at LLM scale.

Three conclusions hold up. Multi-turn RL is not math RL with more steps; the group baseline that made the math recipe cheap is the first thing to break at long horizons, and labs are now choosing between critics and much larger groups. The decisive 2026 advances were systems and objective choices (resumable sandboxes at tens of millions of instances, train–inference alignment, training in the shipping harness, cost in the reward) more than new estimators. And the question that decides the subfield's ceiling, whether RL on agents adds capability rather than sharpening what the base can already do, now has its first agent-specific evidence but still lacks the decisive experiment.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]Z. Hou, Y. Li, J. Tang, Y. Dong et al., “Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning” (SAO), arXiv:2607.07508, 2026-07-08. https://arxiv.org/abs/2607.07508
  2. [2]Z.ai (Zhipu AI), “GLM-5.2: Built for Long-Horizon Tasks,” Z.ai blog, 2026-06-16. https://z.ai/blog/glm-5.2
  3. [3]L. Pu, X. Li, Y. Liu, T. Cao, B. Yang, “Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents” (CANOPY), arXiv:2609.01245, 2026-09-01. https://arxiv.org/abs/2609.01245
  4. [4]Z. Shao et al. (DeepSeek), “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models” (introduces GRPO), arXiv:2402.03300, 2024. https://arxiv.org/abs/2402.03300
  5. [5]M. Chen, L. Wang, B. Qu, “The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents,” arXiv:2608.06663, 2026-08-07. https://arxiv.org/abs/2608.06663
  6. [6]RAGEN team, “RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning,” arXiv:2504.20073, 2025-04. https://arxiv.org/abs/2504.20073
  7. [7]Agentica and Together AI (M. Luo, N. Jain, J. Singh, S. Tan, C. Cai et al.), “DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL,” Together AI blog, 2025-07-02. https://www.together.ai/blog/deepswe
  8. [8]W. Li, S. Im, S. Li (UW–Madison), “Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning,” arXiv:2605.27954, 2026-05-27. https://arxiv.org/abs/2605.27954
  9. [9]L. Tao, B. Peng, W. Yao, T. Ge, H. Cheng, M. H. Wang, J. Gao, S. Li, “TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents,” arXiv:2607.13988, 2026-07-15. https://arxiv.org/abs/2607.13988
  10. [10]J. Yang, Y. Shi, Z. Li, R. Wang, Z. Li, H. Mi, L. Liang, “T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks,” arXiv:2609.11042, 2026-09-10. https://arxiv.org/abs/2609.11042
  11. [11]L. Feng, Z. Xue, T. Liu, B. An, “Group-in-Group Policy Optimization for LLM Agent Training” (GiGPO), NeurIPS 2025, arXiv:2505.10978. https://arxiv.org/abs/2505.10978
  12. [12]S. He, L. Feng, Q. Wei, X. Cheng, L. Feng, B. An, “Hierarchy-of-Groups Policy Optimization” (HGPO), ICLR 2026, arXiv:2602.22817, 2026-02-26. https://arxiv.org/abs/2602.22817
  13. [13]D. Wang, Q. Li, M. Cheng et al. (USTC), “StepPO: Step-Aligned Policy Optimization,” arXiv:2604.18401, 2026-04-20. https://arxiv.org/abs/2604.18401
  14. [14]Cursor, “Composer 2.5,” Cursor blog, 2026-05-18. https://cursor.com/blog/composer-2-5
  15. [15]MiniMax, “MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention,” arXiv:2506.13585, 2025-06. https://arxiv.org/abs/2506.13585
  16. [16]Kimi Team (Moonshot AI), “Kimi K3: Open Frontier Intelligence,” arXiv:2607.24653, 2026-07 (revised 2026-08-07). https://arxiv.org/abs/2607.24653
  17. [17]J. Huang, H. Tang, J. Chen et al. (DeepSeek), “DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale,” arXiv:2609.22978, 2026-09-19. https://arxiv.org/abs/2609.22978
  18. [18]Cognition, “Introducing SWE-2: Pushing the Pareto Frontier,” Cognition blog, 2026-09-10. https://cognition.com/blog/swe-2
  19. [19]M. Marek, M. Ryabinin, “Score Centering Stabilizes Off-policy Reinforcement Learning,” arXiv:2609.20807, 2026-09-17. https://arxiv.org/abs/2609.20807
  20. [20]RadixArk (T. Chen, M. Cheng, S. Dong et al.), “Miles v0.1: Production-Level Post-Training,” arXiv:2609.08368, 2026-09-08. https://arxiv.org/abs/2609.08368
  21. [21]Z. He, S. Zhang, Z. Zhou, Y. Yang, Y. Kang, Y. Zhang, L. K. Qiu, T. Y. Tsui, J. Xu, C. Luo (Microsoft Research), “Agent Lightning v1.0: Towards Harnessed Agentic RL,” arXiv:2608.17528, 2026-08-18. https://arxiv.org/abs/2608.17528
  22. [22]Cursor Research (A. Chan et al.), “Composer 2 Technical Report,” arXiv:2603.24477, 2026-03; blog https://cursor.com/blog/composer-2-technical-report. https://arxiv.org/abs/2603.24477
  23. [23]Cursor, “CursorBench,” Cursor blog, 2026-03-11. https://cursor.com/blog/cursorbench
  24. [24]Cognition (B. Pan, C. Baronio, A. Tam et al.), “Introducing SWE-grep and SWE-grep-mini: RL for Multi-Turn, Fast Context Retrieval,” Cognition blog, 2025-10-16. https://cognition.com/blog/swe-grep
  25. [25]Kimi Team (Moonshot AI), “Kimi K2.5: Visual Agentic Intelligence,” arXiv:2602.02276, 2026-02-02 (revised 2026-08-07). https://arxiv.org/abs/2602.02276
  26. [26]I. Bigio, T. Sanders (OpenAI), “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark,” OpenAI, 2026-07-29. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
  27. [27]Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, S. Song, G. Huang (Tsinghua), “Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?,” NeurIPS 2025 (oral), arXiv:2504.13837. https://arxiv.org/abs/2504.13837
  28. [28]M. Liu, S. Diao, X. Lu, J. Hu, X. Dong, Y. Choi, J. Kautz, Y. Dong (NVIDIA), “ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models,” NeurIPS 2025, arXiv:2505.24864. https://arxiv.org/abs/2505.24864
  29. [29]D. Khatri, L. Madaan et al. (Meta, UT Austin, UCL, UC Berkeley, Harvard), “The Art of Scaling Reinforcement Learning Compute for LLMs” (ScaleRL), ICLR 2026 (oral), arXiv:2510.13786, 2025-10-15. https://arxiv.org/abs/2510.13786
  30. [30]R. Shao et al., “Spurious Rewards: Rethinking Training Signals in RLVR,” arXiv:2506.10947, 2025-06. https://arxiv.org/abs/2506.10947
  31. [31]X. Liu, J. Qian, “Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs,” arXiv:2609.19636, 2026-09-17. https://arxiv.org/abs/2609.19636
  32. [32]M. T. R. Laskar, X.-Y. Fu, S. Bhushan TN, “SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale,” arXiv:2609.17848, 2026-09-15. https://arxiv.org/abs/2609.17848
  33. [33]Cursor, “Real-time RL for Composer,” Cursor blog, 2026-03-26. https://cursor.com/blog/real-time-rl-for-composer
  34. [34]Thinking Machines Lab, “Introducing Inkling,” 2026-07-15, and “Tinker” (https://thinkingmachines.ai/tinker). https://thinkingmachines.ai/news/introducing-inkling/