Research series · 2026 · Report 04 of 16
Report 04 of 16 · Part I · Parametric · Weights

Self-training and distillation

learning from your own traces

The oldest self-improvement loop (generate answers, keep the ones that pass a check, fine-tune on them) still works in 2026, exactly as far as the check can see. The loop's center has moved to on-policy self-distillation, where the student samples its own trajectories and the same model, holding privileged context, grades every token. A burst of July–September 2026 papers shows that the privileged teacher mostly transfers a behavior (confidence, a reasoning mode, the suppression of unlikely tokens) rather than the information it was handed, which is why the loop sharpens what the model already does rather than adding what it lacks (theory says online exploration, unlike filtering, can get past that coverage limit), and why it backfires when the behavior it transfers is "stop searching".

ContentsReport 04 · Parametric
$ tree ./04-self-training-distillation
./04-self-training-distillation
├── 01-a-teacher-that-knows-the-answer# 337 words
├── 02-where-the-loop-came-from# 249 words
├── 03-how-the-loop-works-in-2026# 907 words
├── 04-what-the-privileged-teacher-transfers# 547 words
├── 05-what-the-evidence-shows# 307 words
├── 06-where-it-breaks# 652 words · 1 figure
├── 07-state-of-play-2026# 308 words · 1 figure
├── 08-what-ships# 196 words
└── 09-open-problems# 305 words
9 sections · 43 references · 2 figures
17%
relative accuracy drop from a teacher that sees the answer
59.18 vs 47.48
RL after OPD vs RL after SFT, eight benchmarks
8
prompts that match a 17k-problem dataset
43
references · 21 from Jun 24 – Sep 24, 2026
16 min
reading time
Storage
Weights (full or LoRA)
Engines
On-policy (self-)distillation (per-token reverse KL against a teacher); SFT on self-generated verified trajectories; OPD-then-RL pipelines
Evaluator
Teacher log-probs, often from the same model with a reference solution, skill, or trajectory in context; unit tests; majority votes
Loop timescale
Rounds of hours to days inside a training run; generational across releases
Loop closure
Human on the loop (lab post-training); closed loops only in research prototypes
Evidence maturity
High for verified-trajectory SFT on math and code; medium for on-policy distillation; contested for what privileged self-distillation transfers

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 09

#A teacher that knows the answer

Give a model a math problem and let it reason. Give a second copy of the same model the problem plus the worked solution, and have it score, token by token, how likely it would have been to write what the first copy wrote. Train the first copy to close the gap. The setup is on-policy self-distillation with privileged context (OPSD), and it looks like a free lunch: no reward model, no external teacher, a dense signal at every token, and a grader that cannot be wrong about the final answer because it was handed the final answer 1.

On July 6, 2026, Simran Kaur, Liam Fowl, Sanjeev Arora and colleagues at Princeton reported what happens when the student is a thinking model 2. Across five Qwen3 and OLMo thinking models evaluated on AIME24, AIME25, and HMMT25, privileged-context distillation caused a relative drop of up to 17% in avg@16 accuracy. The damage grew with the amount of privileged context withheld from the student and was worst at long rollout budgets, exactly where thinking models otherwise gain most. The trained models produced fewer verification, backtracking, and hedging markers, even after length normalization. The diagnosis was local: privileged context lowers "fork rates", the frequency of high-entropy positions where several lines of reasoning remain open, and when the student begins a self-correction branch, "privileged OPD penalizes sampled reconsideration tokens that vanilla OPD supports" 2. Distillation from the same teacher without the solution still improved these models, and the privileged variant still helped instruction-tuned models that do little search 2.

The incident compresses the subfield into one failure. Self-training is a loop in which a model learns from its own outputs, filtered or graded by something, and it learns whatever that grader rewards. A grader with the answer in hand rewards confidence rather than the search that produced correct answers. Over the next eleven weeks, at least five groups took OPSD apart and reached a stronger conclusion: much of what the privileged teacher transfers is its behavior, not its information 3,4,5,6,7.

Section 02 / 09

#Where the loop came from

The loop is old. STaR (2022) generated rationales, kept those that reached the correct answer, and fine-tuned on them; when the model failed, it was shown the answer and asked to rationalize backward, the first appearance of answer-conditioned supervision 8. ReST-EM (2023) framed the same filter-and-fine-tune loop as expectation-maximization with binary correctness feedback and beat fine-tuning on human solutions on MATH and APPS, then plateaued after two or three rounds, a shape that recurs across its successors 9. Agent pipelines replaced answer keys with unit tests; SWE-smith (April 2025) trained a 32B model to 40.2% on SWE-bench Verified from 50,000 synthesized tasks 10.

On-policy distillation (OPD) changed the grader's resolution. The student samples its own rollout, a teacher computes log-probabilities on every token, and the student minimizes per-token reverse KL to the teacher: RL's on-policy sampling with distillation's dense signal 11. Thinking Machines' October 2025 write-up took a Qwen3-8B-Base student from 60% to 70% on AIME'24 in about 150 steps, 9 to 30 times cheaper than extrapolated SFT, and quoted the Qwen3 report's comparison of 74.4% for 1,800 GPU hours of OPD against 67.6% for 17,920 GPU hours of RL 11,12. Because the student stays on its own distribution, the blog noted, OPD does not regress "in the self-distillation setting as SFT does" 11. An earlier checkpoint could be the teacher of a later one, and in January 2026 the Self-Distilled Reasoner made the current model its own teacher by putting the gold solution in the teacher's context 1.

Section 03 / 09

#How the loop works in 2026

The teacher is you, with a hint

By mid-2026 OPSD was a field of its own. A survey of recursive self-improvement that coded 1,250 papers (July 2026) counts it as a distinct thread of 56 papers that did not exist 18 months earlier 13. The variants differ mainly in what the teacher is shown. SDFT (January 2026) shows demonstrations and reports less forgetting than SFT 14; Skill-SD (April 2026) shows natural-language skills summarized from completed trajectories and reports gains of 14.0% and 10.9% over GRPO on AppWorld and Sokoban 15; SDAR (May 2026) gates the distillation loss where imperfect skill retrieval makes the teacher unreliable, improving over GRPO by 9.4% on ALFWorld and 10.2% on WebShop 16.

The August wave moved the method into multi-turn agents and ran into a problem math never posed: the student's own actions change the state. SMRC-SD (August 5, 2026) observed that a student acting differently from a reference trajectory reaches states the reference never covered, so the reference stops describing the situation the student is in and becomes "an unreliable source of guidance" 17. It distills only at turns where the student's state matches a state on the successful trajectory, and builds the teacher's context from that matched state. With Qwen3-1.7B, task success rose from 0.746 to 0.865 on ALFWorld and from 0.574 to 0.693 on WebShop against unconditional full-trajectory distillation 17. AgentOPSD (August 6) aggregates token-level teacher–student log-probability gaps into turn-level evidence and updates a Bayesian belief over which turns were pivotal, reaching 89.1% success on ALFWorld with Qwen2.5-7B, ahead of GRPO and other self-distillation baselines 18. LOPD (August) makes the privileged context a learned latent and reports beating RLVR and Skill-SD with under 30% of their rollout budget 19. On September 23, GUI-SD-v2 extended OPSD from GUI grounding to multi-turn phone-agent interaction, first training the model to follow privileged guidance, then selectively distilling step-specific reasoning and memory hints; it reports higher Pass@1 and Pass@3 than existing OPSD baselines on AndroidWorld and MobileWorld 20.

Removing the teacher's information

A second strand asked how little the teacher needs. U-OPSD (August 2026) replaces the gold solution with a majority-vote pseudo-solution from the model's own rollouts and distills on sampled completions whose answers disagreed with that pseudo-solution; averaged over five math benchmarks, it improved Qwen3 non-thinking models by 8.5% at 4B and 10.7% at 8B over the base and beat ground-truth OPSD by 3.2% and 2.3%, while in thinking mode it only matched OPSD 21.

The same question, how little outside supervision a teacher needs, also appears inside RL. RISE (Salesforce, September 2026) builds the teacher from the model's own RLVR trajectory by extrapolating the displacement between the current checkpoint and a trailing anchor, turning a sparse outcome update into a dense token target, and reports gains over RLVR alone and over OPSD across math, code, and agent tasks 22.

The most radical result removes the teacher outright. On August 31, 2026, Yi Ding and Ruqi Zhang measured the supervision a teacher provides during OPD and found substantial noise that grows with teacher scale, to which the student was insensitive: it converged to the same performance whether the noisy supervision was kept or removed 23. Learning concentrated on low-probability tokens, and "using a single fixed negative advantage matches the performance of teacher-provided ones", which led the authors to conclude that "OPD works largely by suppressing low log-probability tokens, which requires no teacher" 23. Their teacher-free replacement, OPSA, raised Qwen3-1.7B's Avg@32 on AIME24 by 35.41 points over the base model and by 16.77 points over OPD 23. A number that large on a 1.7B model needs replication, but the mechanism it points to is precise, and it is sharpening in its plainest form.

Data turned out to matter as little as the teacher. Training OPD on a single query recovers most of full-data OPD's gain; one query's rollouts reach 71.5% of the states full-data training visits, and 16 semantically distinct queries reach 98.9% and match full-data training (September 3) 24. The authors call OPD "data-overfed but algorithm-starved". Data-free OPD (September 12) found that eight prompts match a 17,000-problem dataset, that swapping mathematics for competitive programming "still recovers over ninety percent of the in-domain gain", and that a teacher writing its own training questions matches or beats real data 25. The paper's reading is that "OPD transfers the teacher's mode of reasoning rather than knowledge related to the data" 25.

Self-training on verified agent trajectories

The filter-and-fine-tune loop did not disappear; it moved inside agent RL pipelines. Socratic-SWE (June 2026) distills its own solving traces into skills describing recurring failures, uses them to generate targeted repair tasks in real repositories, validates the tasks by execution, and trains on successful repairs; after three rounds it reached 50.40% on SWE-bench Verified, ahead of other self-evolving baselines under the same compute 26. A September 20 paper on software-engineering agents found that pooled RL across task categories produces a "category see-saw", gains in some categories and regressions in others 27. Its fix trains category experts that refresh which instances they have mastered, reuse "its own verified successful trajectories for Repair SFT", and reselect tasks for more RL, then consolidates the experts into one student with multi-teacher OPD. No external model supplies solutions; the final policy resolves 58.04% on the Pro-618 split and 59.00% on SWE-bench Multilingual, 5.39 and 2.78 points over its base 27. In both systems the acceptor is execution: training data counts only if the tests pass.

Section 04 / 09

#What the privileged teacher transfers

If the teacher's noise can be dropped and its data can be eight prompts, what does the privileged context add? Five papers between August 10 and September 22, 2026 tested that directly, and they converge.

Ichihara and colleagues (August 10) replaced the paired reference solution with a problem and solution from a different example, keeping everything else fixed. Across three models and three math benchmarks, this "OP²SD" improved over the base model and stayed competitive with OPSD, implying "that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor" 3. Shrestha and Tessier (August 18) found with Qwen3 models from 1.7B to 8B that the correct reference "does not provide a consistent performance benefit", that a solution from another problem can outperform the correct one on several benchmarks, and that the student's predictions align more with the base model's thinking behavior than with anything the reference supplies 4. Zhang and colleagues (September 17) built AMPLE-Math, 5,319 problems with six reasoning views sharing each answer, and compared each view with matched reference-free distillation: reference-free distillation accounted for much of Qwen3-1.7B's improvement, complete traces added about two points for SmolLM3-3B at step 50, and switching the student from short direct-response rollouts to long thinking rollouts turned gains into losses in both model families 5. Their conclusion is that a reference's value "is what it adds to this cross-mode transfer, not how much of the solution it reveals" 5.

Two papers measured the behavior itself. Baumann, Hübotter, Krause and colleagues (September 18) contrasted "attractive" self-distillation toward a privileged teacher with "repulsive" distillation away from one: attraction "suppresses exploratory reasoning and promotes shorter, more confident responses", repulsion lengthens responses and destabilizes, and combining attraction to a correct-solution teacher with repulsion from an incorrect-solution teacher cancels the shared behavioral shift, leaving a signal that tracks correctness with stable response lengths 6. That is the Princeton backfire's mechanism, isolated and reversed. Tian and colleagues (September 22) varied what a self-teacher sees and found that abstractions (a named strategy, a framing, a problem category) beat the full solution by 1.4 points at 4B and 1.6 at 8B while storing an order of magnitude fewer hint tokens; answer-only conditioning stayed within 0.2 points of the full solution, and initial teacher–student KL did not predict downstream performance 7. "What a self-teacher should see," they conclude, "is therefore not everything it could, but the level of abstraction its student can still act on" 7.

The picture is consistent across groups. The privileged context changes how the teacher behaves (more confident, shorter, in a different reasoning mode), and the student mostly learns that behavior. When the behavior is helpful, as with a direct-response model borrowing a thinking model's habits, OPSD works even with the wrong reference. When the behavior is "skip the check", as with a thinking model on long rollouts, it degrades the student regardless of how correct the reference is. An August 26 review of the literature names the resulting failure "collapse, the progressive narrowing of the set of reasoning paths the model can produce", says privileged information aggravates it, and organizes the fixes around three levers: which tokens get weight, what the teacher is shown, and how the guidance changes over training 28.

Section 05 / 09

#What the evidence shows

Date System Setting Result Baseline / context
Sep 23, 2026 RL Starts Before RL 29 8 text benchmarks, average after RL OPD→RL 59.18 SFT→RL 47.48; Base→RL 51.90
Sep 22, 2026 Privileged context design 7 Competition math, 4B / 8B Abstractions +1.4 / +1.6 pts over full solution Answer-only within 0.2 pts of full solution
Sep 20, 2026 Category-aware SWE experts 27 SWE-bench Pro-618; Multilingual 58.04% (+5.39); 59.00% (+2.78) Base model; no external solution trajectories
Sep 12, 2026 Data-free OPD 25 Math, multi-teacher 8 prompts match 17k; 1k self-written questions close 98.5% of headroom 96.6% with 7k real examples
Sep 3, 2026 One-example OPD 24 State coverage 1 query 71.5%; 16 queries 98.9%, match full data Full-data OPD
Aug 31, 2026 OPSA, teacher-free 23 AIME24 Avg@32, Qwen3-1.7B +35.41 pts over base; +16.77 over OPD Fixed negative advantage matches teacher signal
Aug 2026 U-OPSD 21 5 math benchmarks, Qwen3-4B non-thinking +8.5% avg over base; +3.2% over OPSD Parity with OPSD in thinking mode
Aug 6, 2026 AgentOPSD 18 ALFWorld, Qwen2.5-7B 89.1% success GRPO and self-distillation baselines
Aug 5, 2026 SMRC-SD 17 ALFWorld / WebShop, Qwen3-1.7B 0.746→0.865 / 0.574→0.693 Unconditional full-trajectory distillation
Jul 6, 2026 Privileged OPSD, thinking models 2 AIME24/25, HMMT25; 5 models Up to 17% relative drop, avg@16 Before distillation
Jun 2026 Socratic-SWE 26 SWE-bench Verified 50.40% after 3 rounds Self-evolving baselines, same compute
May 2026 SDAR 16 ALFWorld / WebShop +9.4% / +10.2% GRPO alone
Oct 2025 OPD, Thinking Machines 11 AIME'24, Qwen3-8B-Base 60% → 70% in ~150 steps 9–30× cheaper than extrapolated SFT

The margins are real and conditional. Most 2026 rows beat GRPO or vanilla OPD on specific benchmarks in a single group's setup, and none has an independent replication yet. The rows that carry the most information are the ablations: the ones showing that the teacher's information, its noise, and its data can each be removed with little loss, because they say what the loop runs on.

Section 06 / 09

#Where it breaks

A grader that knows the answer suppresses search

The mechanism behind the July backfire applies well beyond math 2. Reverse-KL distillation pulls the student toward the teacher at every token the student generates. At a forking position, where a thinking model might continue or stop to verify, the unconditioned model spreads probability across both. A teacher conditioned on the answer has no reason to verify, so it puts low probability on the checking branch, and the student is penalized each time it forks. At test time the answer is gone and the confidence remains. The attractive-teacher result shows the same shift without thinking models in the picture: moving toward a privileged teacher makes responses shorter and more confident as a behavior in its own right 6.

Fig 04.1 · A teacher with the answer stops searchPrinceton · thinking models · 2026
ONE FORK IN A STUDENT ROLLOUT VANILLA OPD teacher: no solution PRIVILEGED OPD teacher sees the solution Student problem only FORK Continue toward the answer Reconsider verify, backtrack, hedge supported supported supported penalized AFTER PRIVILEGED OPD, THINKING MODELS fewer verification, backtracking and hedging markers; lower fork rates RELATIVE DROP IN AVG@16 up to 17% AIME24, AIME25, HMMT25 5 Qwen3 and OLMo models
The penalty lands on the check. The student samples its own rollout and the same model grades every token. Without the solution in context the teacher supports both branches at a fork; with it, the teacher puts low probability on reconsideration, so reverse-KL training penalizes each sampled check. On five thinking models the result was fewer verification markers and a relative avg@16 drop of up to 17% 2.

Agent settings add a second failure. Whatever the teacher knows that the student will not know at inference, the student learns to act as if it knew, and in multi-turn tasks the teacher's knowledge can also be about the wrong state; SMRC-SD's gains came from refusing to distill where the reference no longer described the student's situation 17. Earlier in 2026, Li et al. found that OPD transfers little from a teacher whose thinking patterns the student cannot reach: where OPD worked, alignment concentrated on a shared set of high-probability tokens carrying 97% to 99% of the probability mass at student-visited states 30. Teachers should be chosen by distributional closeness, not leaderboard rank. Mundane failures surface too: a September 17 study traced runaway response length in OPD to base students and post-trained teachers placing their stopping probability on different end-of-sequence tokens, even with identical declared stop sets 31.

What theory says the loop can and cannot do

The 2026 results fit a theory published in December 2024. Huang et al. showed that self-improvement cannot "create information that is not already in the model"; it sharpens, moving probability onto good sequences the model already generates. SFT on filtered samples is minimax optimal when the model already covers good answers, and an RL-style approach can improve on it by exploring online, bypassing the coverage requirement 32. The accurate reading is narrower than "self-training cannot create new capability": filter-and-fine-tune loops are capped by what the model already samples, and loops that explore can pass that cap, though not the information available to their verifier. A related result found that the generation–verification gap that bounds self-improvement grows with pretraining compute 33.

September 2026 supplied the empirical version. "Sequential Beats Joint" (September 3) found that OPD followed by RL beats pure OPD, pure RLVR, and every tested way of fusing the two signals in one update, and explained it through pass@k and parameter updates: "OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support", while optimizing both at once makes them interfere 34. OPD also made a better RL cold start than SFT, and its validation score was the signal for when to switch 34. OPSA's finding that OPD's gains come from suppressing low-probability tokens is sharpening described at the token level 23. The loop's ceiling is set by what the teacher's support and the verifier can see; OPD's sampling decides how efficiently the student reaches that support, then RL sharpens inside what the verifier can identify.

Collapse and the search alternative

The collapse the August review describes is a within-model narrowing of reasoning paths, not the classic model collapse from training generations of models on each other's outputs, which accumulating real data prevents 28,35. Both share a cause: a signal that rewards the model for being more like its most confident self. Weight-space search does not escape the rule. Evolution strategies now match GRPO's accuracy with larger, off-task weight changes (COLM 2026) and, per an August 2026 study, avoid GRPO's entropy collapse and reach higher Pass@K 36,37. They replace the proposer, not the acceptor: the fitness function plays the filter's role and has its blind spots.

Section 07 / 09

#State of play, 2026

Three things changed between June and September 2026.

The field's understanding of OPSD inverted. In the first half of the year, privileged context was treated as information the teacher passed down, and variants competed on richer context. By late September, five independent ablations had shown that wrong references, abstractions, or answer-only context often work as well, and that the damage and the benefit both come through behavior 3,4,5,6,7. The practical rules that follow: keep answer-level context away from thinking models on long rollouts, distill only where the teacher's context describes the student's state, and prefer abstractions the student can act on.

OPD's role shifted from replacement to preparation. "RL Starts Before RL" (September 23) found that OPD-initialized students reach higher performance after RL than direct RL or SFT-then-RL, even "when OPD produces little immediate improvement in accuracy": on the text setting, OPD then RL averaged 59.18 across eight benchmarks, against 47.48 for SFT then RL and 51.90 for RL from the base, winning on all eight 29. Together with "Sequential Beats Joint", two groups in one month converged on OPD first, RL second 29,34. Both are preprints; if they hold, pre-RL accuracy is the wrong acceptor for the self-training stage.

Fig 04.1 · OPD first, RL secondRL Starts Before RL · Sep 2026
STAGE BEFORE RL AVERAGE SCORE AFTER RL, 8 TEXT BENCHMARKS 0 None (base model) 51.90 SFT 47.48 OPD 59.18 wins on all eight
Distill first, then reinforce. Final averages across eight text benchmarks after RL, starting from the base model, an SFT checkpoint, or an OPD checkpoint; OPD then RL wins on all eight, even when OPD alone produces little immediate improvement in accuracy 29. A separate study reads the ordering as OPD expanding coverage and RL sharpening within that support 34.

The context-to-weights question sharpened, and on-policy sampling did not always win it. Experience Distillation (July 2026) distilled an agent's accumulated experience into the same checkpoint without that context: teacher-sampled forward-KL distillation kept 64.8% of the in-context gain on 749 SWE tasks, plain SFT kept 3.8%, and on-policy reverse KL at matched compute kept 9.1% and 0.4% on two text games, because the student without the experience rarely samples the trajectories where it helps 38. Report 12 (Consolidation and co-evolution) owns that question. The relevant point here is the coverage condition again: the student's own rollouts are the right training distribution only when the improvement is reachable from them.

Section 08 / 09

#What ships

On-policy distillation now appears in shipped post-training, not just in papers. SEA-LION v4.8 (September 2026), built on NVIDIA Nemotron 3, was post-trained "with supervised fine-tuning and online on-policy distillation", raising its 120B model's SEA-HELM score from 49.30 to 63.44 39. The KwaiMind image editor (September 2026) trains specialized policies against separate rewards, "which are consolidated through on-policy distillation" 40. The Data-free OPD authors describe OPD as "a standard component of frontier post-training pipelines" 25.

Self-training on deployment traffic remains rare and early. NeoHorse-1 (September 8) records, for each user turn, the predicted capability demand, the model tier its router chose, and the interaction that followed, converts validated records into training examples, and adds routing-guided OPD; post-training lifted an eleven-benchmark macro-average from 58.94 to 64.87 at 4B 41. The closest production analogue to self-distillation remains Cursor's Composer 2.5 (May 2026), which inserts feedback at the point in a trajectory where the model could have done better and uses "the resulting model distribution as a teacher" 42; report 05 covers Cursor's loop. Weight-level self-training after deployment, driven by an agent's own traces, is still the exception, consistent with the series' finding that shipped self-improvement mostly lives in text.

Section 09 / 09

#Open problems

What the teacher should see. The August–September ablations show privileged context works mostly through induced behavior, and that the best context depends on student scale and task 7. Nobody can yet predict, before training, which context will teach a useful behavior and which will teach overconfidence; initial teacher–student KL does not 7. Nearly every self-distillation method depends on that choice.

When to stop distilling and start RL. If OPD's value is the coverage it prepares for RL, the self-training stage needs an acceptor that measures RL-readiness rather than accuracy 29,34. "Sequential Beats Joint" suggests OPD validation score; a general answer would change how labs schedule post-training.

Self-training where no filter exists. Every strong result above has an answer key, a test suite, a deterministic environment, or a stronger checkpoint. Open-ended agent work has none, and the substitutes (majority votes, learned judges, self-written skills) inherit their grader's errors. Report 03's verifier problem is this report's ceiling.

Repeated self-writes into weights. SEAL (2025) showed a model can write its own fine-tuning data but forgot earlier tasks as edits accumulated 43; on-policy distillation reduces forgetting 11,14. Nobody has shown a self-directed loop writing into weights hundreds of times without drift, which is the gap between generational self-training in a lab and a model that improves from its own deployment.

A model learning from its own traces improves as far as its grader can see: to the edge of its coverage when the grader is a filter, somewhat further when the loop explores, and in the wrong direction when the grader's confidence comes from knowledge the student will not have. The 2026 work adds a harder lesson. The dense teacher transfers what it does more than what it knows, so the question for any self-distillation loop is what behavior the teacher's context induces, not how much the context contains.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]S. Zhao, Z. Xie, M. Liu, J. Huang, G. Pang, F. Chen, A. Grover (Meta, UCLA), Self-Distilled Reasoner (on-policy self-distillation, OPSD), arXiv 2601.18734, January 2026 (rev. March 2026). https://arxiv.org/abs/2601.18734
  2. [2]S. Kaur, N. Ri, Y. He, L. Fowl, S. Arora (Princeton), “Rethinking On-Policy Self-Distillation for Thinking Models,” arXiv 2607.05184, July 6, 2026. https://arxiv.org/abs/2607.05184
  3. [3]Y. Ichihara, N. Iwase, M. A. Quamar, J. Komiyama, “Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation,” arXiv 2608.09228, August 10, 2026. https://arxiv.org/abs/2608.09228
  4. [4]S. Shrestha, A. Tessier, “Rethinking Privileged Information in On-Policy Self-Distillation,” arXiv 2608.18271, August 18, 2026. https://arxiv.org/abs/2608.18271
  5. [5]X. Zhang, W. Chow, J. Fang, X. Zhu, Z. Liang, T.-S. Chua, “What Does Privileged Information Add to On-Policy Self-Distillation?,” arXiv 2609.20612, September 17, 2026. https://arxiv.org/abs/2609.20612
  6. [6]A. Baumann, A. Ashirmatov, L. Schmidt-Traub, F. Lübeck, J. Hübotter, T. K. Buening, A. Krause, “On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation,” arXiv 2609.21561, September 18, 2026. https://arxiv.org/abs/2609.21561
  7. [7]K. Tian, S. Liu, T. Jiang, S. Dong et al., “What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation,” arXiv 2609.25623, September 22, 2026. https://arxiv.org/abs/2609.25623
  8. [8]E. Zelikman, Y. Wu, J. Mu, N. D. Goodman, “STaR: Bootstrapping Reasoning With Reasoning,” arXiv 2203.14465, March 2022 (NeurIPS 2022). https://arxiv.org/abs/2203.14465
  9. [9]A. Singh, J. D. Co-Reyes, R. Agarwal et al. (Google DeepMind), “Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models” (ReST-EM), arXiv 2312.06585, December 2023; TMLR 2024. https://arxiv.org/abs/2312.06585
  10. [10]J. Yang, C. E. Jimenez et al. (Princeton), “SWE-smith: Scaling Data for Software Engineering Agents,” arXiv 2504.21798, April 2025. https://arxiv.org/abs/2504.21798
  11. [11]K. Lu and Thinking Machines Lab, “On-Policy Distillation,” Thinking Machines blog, October 27, 2025. https://thinkingmachines.ai/blog/on-policy-distillation
  12. [12]Qwen Team (A. Yang et al.), “Qwen3 Technical Report,” arXiv 2505.09388, May 2025. https://arxiv.org/abs/2505.09388
  13. [13]M. Chen, L. Wang, B. Qu, “Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops,” arXiv 2607.07663, July 8, 2026 (v2 September 2026). https://arxiv.org/abs/2607.07663
  14. [14]I. Shenfeld et al., SDFT (self-distillation fine-tuning for continual learning), arXiv 2601.19897, January 2026 (rev. August 2026). https://arxiv.org/abs/2601.19897
  15. [15]H. Xiao, H. Wang et al., Skill-SD (skill-conditioned self-distillation), arXiv 2604.10674, April 2026. https://arxiv.org/abs/2604.10674
  16. [16]Z. Lu, Z. Yao et al., SDAR (self-distilled agentic reinforcement learning), arXiv 2605.15155, May 2026. https://arxiv.org/abs/2605.15155
  17. [17]J. Liu, W. Li, J. Ling, P. Wang, “When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents” (SMRC-SD), arXiv 2608.05219, August 5, 2026. https://arxiv.org/abs/2608.05219
  18. [18]Z.-H. Wang, Z. Lu, Z. Yao, J. Wu et al., “AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning,” arXiv 2608.05987, August 6, 2026. https://arxiv.org/abs/2608.05987
  19. [19]G. Zhang et al., “Latent On-Policy Self-Distillation” (LOPD), arXiv 2608.13040, August 2026. https://arxiv.org/abs/2608.13040
  20. [20]Y. Zhang, D. Wu, H. Shen et al., “Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents” (GUI-SD-v2), arXiv 2609.27307, September 23, 2026. https://arxiv.org/abs/2609.27307
  21. [21]Y. Li, B. Wang, Y. Liang, Y. Tian, D. Fu, N. Vasconcelos, “On-Policy Self-Distillation without Any Supervision” (U-OPSD), arXiv 2608.06296, August 2026. https://arxiv.org/abs/2608.06296
  22. [22]Y. Li, S. Yavuz, S. Joty (Salesforce AI Research), “RISE: Recursive Improvement via Self-Extrapolating Policy Distillation,” arXiv 2609.05295, September 2026. https://arxiv.org/abs/2609.05295
  23. [23]Y. Ding, R. Zhang, “Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement” (OPSA), arXiv 2608.31046, August 31, 2026. https://arxiv.org/abs/2608.31046
  24. [24]Z. Fu, B. He, Y. Zuo et al., “Rethinking On-Policy Distillation of Large Language Models II: One Training Example,” arXiv 2609.04172, September 3, 2026. https://arxiv.org/abs/2609.04172
  25. [25]G. Li, M. Zheng, M. Song et al., “Data-free On-policy Distillation,” arXiv 2609.14193, September 12, 2026 (v2 September 17). https://arxiv.org/abs/2609.14193
  26. [26]C. Xiao, Z. Jiao, S. Wang, W. Zhang, L. Qu, Socratic-SWE (self-evolving SWE agents via trace-derived skills), arXiv 2606.07412, June 2026 (under review). https://arxiv.org/abs/2606.07412
  27. [27]J. Zhao, Z. Jiang, S. Zheng, M. Shan, X. Xu, L. Qu, “One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents,” arXiv 2609.23377, September 20, 2026. https://arxiv.org/abs/2609.23377
  28. [28]J. Robert, R. Qader, “One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation,” arXiv 2608.25936, August 26, 2026. https://arxiv.org/abs/2608.25936
  29. [29]S. Dong, Y. Zhu, Y. Xu et al., “RL Starts before RL: On Policy Distillation for Better Reinforcement Learning,” arXiv 2609.28145, September 23, 2026. https://arxiv.org/abs/2609.28145
  30. [30]Y. Li, Y. Zuo et al. (Tsinghua / THUNLP), Rethinking on-policy distillation, arXiv 2604.13016, April 2026. https://arxiv.org/abs/2604.13016
  31. [31]Y. Yang, T. Yu, S. Li et al., “When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation,” arXiv 2609.20511, September 17, 2026. https://arxiv.org/abs/2609.20511
  32. [32]A. Huang, A. Block, D. J. Foster, C. Zhang, M. Simchowitz, “Self-Improvement in Language Models: The Sharpening Mechanism,” arXiv 2412.01951, December 2024. https://arxiv.org/abs/2412.01951
  33. [33]Y. Song, H. Zhang, C. Eisenach et al., “Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models,” arXiv 2412.02674, December 2024. https://arxiv.org/abs/2412.02674
  34. [34]B. Li, B. Chen, C. Yang, P. Nie, C. Zhao, X. Ye, “Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR,” arXiv 2609.04108, September 3, 2026. https://arxiv.org/abs/2609.04108
  35. [35]M. Gerstgrasser, R. Schaeffer, A. Dey et al., “Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data,” arXiv 2404.01413, April 2024 (COLM 2024). https://arxiv.org/abs/2404.01413
  36. [36]W. Hoy, B. Wang, X. Pan, Matching accuracy, different geometry: ES vs GRPO in LLM post-training, arXiv 2604.01499, April 2026 (COLM 2026). https://arxiv.org/abs/2604.01499
  37. [37]Y. Ba, Z. Zheng, Y. Xie et al., Understanding ES for LLM reasoning: broader coverage than GRPO, arXiv 2608.27351, August 2026. https://arxiv.org/abs/2608.27351
  38. [38]C. Gou, H. Tu, Y. Fang, J. Cai, H. Rezatofighi, Sample-efficient learning from agent experience (Experience Distillation), arXiv 2607.21051, July 2026. https://arxiv.org/abs/2607.21051
  39. [39]A. Aulia, A. Dabeer, A. Jeongmi et al., “SEA-LION-v4.8: A Technical Report,” arXiv 2609.18310, September 16, 2026. https://arxiv.org/abs/2609.18310
  40. [40]J. Wu, Z. Li, Y. Hu et al., “KwaiMind Technical Report,” arXiv 2609.26375, September 22, 2026. https://arxiv.org/abs/2609.26375
  41. [41]NeoHorse Team (G. Cao, G. Dai, T. Guo, K. Han et al.), “NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness,” arXiv 2609.08183, September 8, 2026. https://arxiv.org/abs/2609.08183
  42. [42]Cursor, “Composer 2.5,” Cursor blog, May 18, 2026. https://cursor.com/blog/composer-2-5
  43. [43]A. Zweiger, J. Pari, H. Guo, E. Akyürek, Y. Kim, P. Agrawal (MIT), “Self-Adapting Language Models” (SEAL), arXiv 2506.10943, June 2025. https://arxiv.org/abs/2506.10943