Research series · 2026 · Report 03 of 16
Report 03 of 16 · Part I · Parametric · Weights

Self-generated rewards

when the model grades its own work

When no ground truth exists, a model can supply its own training signal by judging itself, trusting its own confidence or consensus, or writing rubrics. Used without an outside anchor, all three mostly sharpen what the model already knows: they rise and then fall within a few iterations, they can be gamed (a judge shown a candidate answer scores its plausibility, not its correctness), and they are bounded by the gap between what a model can generate and what it can verify. The 2026 work that holds up anchors the signal to something the policy cannot talk past: a judge that solves before it reads, decorrelated peers, execution, hidden verifiers, and rubrics checked against a stronger reference.

ContentsReport 03 · Parametric
$ tree ./03-self-generated-rewards
./03-self-generated-rewards
├── 01-more-convincing-not-more-correct# 361 words · 1 figure
├── 02-why-the-reward-is-the-hard-part# 287 words
├── 03-judging-yourself# 187 words
├── 04-trusting-your-own-confidence# 641 words
├── 05-rubrics-for-tasks-without-answers# 405 words
├── 06-verifying-agents# 464 words
├── 07-what-the-evidence-shows# 407 words
├── 08-where-it-breaks# 690 words · 1 figure
├── 09-the-ceiling-the-solver-verifier-gap# 241 words
├── 10-state-of-play-2026# 220 words
├── 11-what-ships# 176 words
└── 12-open-problems# 319 words
12 sections · 34 references · 2 figures
0.72 → 0.94
judge pass rate while true accuracy stays at 0.20
0.719 → 0.012
false positives once the judge commits to its own answer
57.2%
GLM 5.2 rollouts that hack on DeepSWE
34
references · 11 from Jun 24 – Sep 24, 2026
18 min
reading time
Storage
Weights (policy and, often, the judge/verifier itself)
Engines
Gradient updates (GRPO, DPO) driven by self-judgment, confidence, consensus, rubrics, or trained verifiers
Evaluator
Self-judgment and intrinsic signals at the bottom; rubric judges and trained critics in the middle; execution, hidden verifiers and production outcomes at the top
Loop timescale
Offline training runs (hours to days per iteration); test-time RL on unlabeled batches
Loop closure
Human on the loop (lab post-training)
Evidence maturity
Dense stream of 2026 preprints, several ICLR/ACL 2026 papers, two frontier-lab disclosures in July–August 2026 (Kimi K3, Anthropic); few matched-budget or cross-family replications

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 12

#More convincing, not more correct

In July 2026 a single-author preprint put a number on the oldest worry about self-rewarding. Chenyu Zhou trained Qwen3 policies on GSM8K against a reference-free LLM judge (one that sees the question and the candidate answer, but no reference solution), the setup behind self-rewarding, self-play and most LLM-as-a-judge pipelines. Over training, the judge's pass rate climbed from 0.72 to 0.94. True accuracy, measured by a held-out exact-match check the judge never saw, stayed at 0.20 across three seeds 1. The policy had not learned arithmetic. It had learned to write answers the judge found convincing.

The paper's diagnosis is structural rather than a bug in one model: "conditioned on a candidate, a judge scores plausibility, not correctness, leaving false-positive basins a policy learns to exploit" 1. The fooling answers transferred across judge families (Qwen, Llama, Gemma) and scales, and a strict three-judge ensemble still accepted 55% of them 1. More judges did not help, because every judge was asked the same wrong question.

One change did help. When the judge committed to an answer of its own before looking at the candidate, the false-positive rate fell from 0.719 to 0.012; used as the training reward, the de-anchored judge kept false positives at zero 1. The paper also offers a falsifiable bound: the gap between judged and true accuracy is at most one minus the policy's accuracy 1, so weak policies on hard tasks have the most room to fool their graders. The work is an unreviewed preprint on grade-school math, and its bound has not been tested elsewhere. But it compresses the field's lesson into one experiment: a self-generated reward works until the policy finds where the grader's judgment and the truth come apart, and the fix is to change what the grader is anchored to, not to add more graders.

Fig 03.1 · Convincing the judge, not being rightZhou · Qwen3 on GSM8K · Jul 2026
SELF-PLAY AGAINST A REFERENCE-FREE JUDGE 1 0 START END OF TRAINING 0.72 0.94 judge pass rate 0.20 0.20 true accuracy held-out exact match FALSE-POSITIVE RATE Judge reads the candidate 0.719 Judge commits to its own answer first 0.012 0 1
The judge is fooled, not the metric. Left: Qwen3 policies trained against a reference-free judge raise its pass rate from 0.72 to 0.94 while true accuracy on a held-out exact-match check stays at 0.20 across three seeds; only the start and end values are reported, so the dashed line marks endpoints, not a trajectory. Right: making the judge commit to its own answer before reading the candidate cuts its false-positive rate from 0.719 to 0.012 1.

That pattern runs through the 2026 literature on self-generated rewards. They work, measurably, and mostly by sharpening (concentrating probability on answers the model could already produce). They rise and then fall. Their failures trace back to the evaluator rather than the proposer. The systems that keep improving tie the self-generated signal to something the model cannot talk its way past.

Section 02 / 12

#Why the reward is the hard part

The series asks five questions of every self-improving system: where the change is written, what proposes it, what decides it is better, who signs off, and whether gains beat matched compute. Here the first two answers are fixed: the change is written into weights, and a gradient step proposes it. Everything interesting happens at the third. The evaluator (the component that produces the score an update is judged by) sets a ceiling on the loop, and the available evaluators fall into a rough hierarchy: formal checks above execution and tests, above learned judges, above self-judgment, above intrinsic signals such as confidence.

Math and code sit near the top, which is why agentic RL leans on them (see 01, Agentic RL). Most of what people want from agents sits lower: a research summary, a support reply, a refactor no test pins down, a web task whose success is a judgment call. Human labels do not scale to the volume RL consumes. The pull toward letting the model grade itself is economic before it is scientific.

The scientific case rests on one asymmetry: checking is often easier than doing. If a model recognizes good answers more reliably than it produces them, it can sample many, keep the ones it rates highly, and train toward them. The gap between verification and generation ability is the headroom, and every method below is a bet on where that gap lives and how to use it without the model learning to exploit the grader instead. The naive version fails for predictable reasons: a self-judge inherits the policy's blind spots and biases, and because judge and policy often share weights, the policy drifts toward whatever the judge likes while the judge drifts with it.

Section 03 / 12

#Judging yourself

The self-rewarding lineage is short. Meta's Self-Rewarding Language Models (January 2024) had one Llama-2-70B model generate responses, score them with an LLM-as-a-judge prompt, and train on the resulting preference pairs, improving over three iterations 2. Meta-Rewarding (July 2024) added a meta-judge that grades the model's own judgments, lifting Llama-3-8B-Instruct's AlpacaEval 2 win rate from 22.9% to 39.4% 3. CREAM (ICLR 2025) explained the stall that followed: a self-judge assigns confident scores to pairs it cannot distinguish, and the noise compounds across iterations; down-weighting pairs whose ranking flips between iterations extended the useful run 4.

Every member of the lineage plateaus within a few iterations, and the July 2026 plausibility result explains why more precisely than the lineage itself did. A judge shown a candidate is not verifying it; it is rating how much the candidate looks like a correct answer. Meta-judging and consistency filtering improve that rating. They do not change what is being rated. Zhou's commit-first judge does, and so does every 2026 method below that holds up: each replaces "does this look right?" with a question whose answer does not depend on the candidate's persuasiveness.

Section 04 / 12

#Trusting your own confidence

A second branch dropped the judge prompt and asked whether the model's own distribution already contains a usable reward: self-certainty over tokens, entropy over answers, or agreement across samples. The field calls this reinforcement learning from internal feedback, or unsupervised RLVR. Two 2025 results set the expectations. INTUITOR used self-certainty as its sole reward and matched GRPO on math, though a frozen scorer was exploited "around the 100th update step" when the policy learned to append an already-solved problem to its answers 5. TTRL rewarded agreement with the majority vote on unlabeled test questions and lifted Qwen-2.5-Math-7B's AIME 2024 pass@1 by "approximately 211%" 6.

The March 2026 answer to what those gains are made of came from a group that included TTRL's first author. "How Far Can Unsupervised RLVR Scale LLM Training?" (ICLR 2026) splits label-free methods into intrinsic ones (rewards from the model's own confidence or consensus) and external ones, and proves that "all intrinsic methods converge toward sharpening the model's initial distribution" 7. Sharpening "succeeds when initial confidence aligns with correctness but fails catastrophically when misaligned." Across methods, intrinsic rewards "consistently follow a rise-then-fall pattern," with "collapse timing determined by model prior rather than engineering choices" 7. The authors propose a Model Collapse Step as a measure of that prior and a practical indicator of whether a model is trainable this way, and they find intrinsic rewards still useful for test-time training on small datasets 7.

The theory also explains the strangest result in the area. In June 2025, random rewards raised Qwen2.5-Math-7B by 21.4 points on MATH-500, close to the 29.1 points from ground-truth rewards, but "often fail to produce gains for other model families, such as Llama3 or OLMo2" 8. If intrinsic rewards only amplify what the model already prefers, then on a model whose preferences are already mostly right, almost any reward that triggers amplification looks like a good reward. Any new self-generated reward has to beat a random or format-only baseline on the target model family before it counts as a reward at all.

The consensus trap and how 2026 escapes it

Majority-vote rewards fail in a specific way. CoVerRL (ACL 2026) named it: "as training maximizes self-consistency, output diversity collapses, causing the model to confidently reinforce systematic errors that evade detection. We term this the consensus trap" 9. Its fix alternates one model between generator and verifier roles, and it "outperforms label-free baselines by 4.7–5.9% on mathematical reasoning benchmarks" 9. TTRL-CoCoV (June 2026) routes low-confidence samples through a verifier and adds an exploration bonus to high-confidence ones, reporting +9.8% pass@1 and +18.7% pass@16 over TTRL 10.

August 2026 brought two sharper diagnoses. OM-GRPO locates the collapse in the gradient: it "arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning" 11. Masking gradients on the answer span while keeping a soft, frequency-based answer reward, the method matches supervised ground-truth training across three backbones and beats majority voting by 4.24 points in the test-time training setting 11. Co-RL attacks the correlation directly. It trains several models that share no parameters, each rewarded by its peers, and shows that "increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops" 12. It reports average gains of 3.0–8.6% across seven text benchmarks without ground-truth labels, matching or beating supervised methods 12.

The escape, in every case, is a second signal whose errors differ from the first. CoVerRL's verifier is the same model, so its gains are real and bounded. OM-GRPO stops the consensus signal from writing itself directly into the answer tokens. Co-RL buys independence with different models. None of them removes the ceiling the March 2026 theory describes; they delay the fall and raise the peak.

Section 05 / 12

#Rubrics for tasks without answers

The methods above live in math because math has answers to vote on. For open-ended tasks, the 2025 response was to decompose "is this good?" into narrow criteria and reward the checklist: Scale AI's Rubrics as Rewards reported gains of up to 31% on HealthBench over a Likert-scale judge 13, and RLCF was the only method among those tested to improve all five instruction-following benchmarks on Qwen2.5-7B 14. Narrow questions ("does the response state the dosage?") sit closer to verification than holistic ones.

Moonshot AI's Kimi K3 technical report (July 2026) shows what that idea looks like at 2.8T parameters 15. For non-verifiable tasks K3 uses an "Agentic Generative Reward Model" that compares candidates in tournament-style binary matchups and must follow a mandatory protocol: "(1) read the outcome, product, or text output; (2) generate a rubric; (3) score each candidate against the rubric; and (4) record the rubric-assigned scores in a scorepad" 15. The judge writes its criteria before it scores, the same move as Zhou's commit-first judge, applied to rubrics. K3 also hard-codes a defense against a known hack: "to mitigate reward hacking toward increasingly verbose outputs," any candidate longer than a multiple of the cold-start model's typical length "automatically loses the binary comparison" 15. Its predecessor K2 (July 2025) had already anchored its self-critique rubric reward by refining the critic "using verifiable signals" from on-policy rollouts 16.

The July–September 2026 papers treat the rubric reward itself as the object to fix. CriPO (July 2026) found a failure that has nothing to do with the judge: criteria that some rollouts satisfy but whose signal is erased because scalar aggregation gives those rollouts non-positive advantage. It calls them Suppressed Criteria and finds them in "over 57% of samples" throughout training, 1.8 per sample on average; its fix uses a counterfactual self-teacher to find the criterion-relevant tokens in negative-advantage rollouts and flips their advantages to positive 17. RLVR² (September 2026) converts each criterion's scores into within-group rankings rather than summing them, and reports controlling formatting and reasoning-length effects across 16 benchmarks 18. And an August 2026 study (EMNLP 2026 Findings) found that a Qwen3-1.7B probe judge, reading criterion satisfaction from activations rather than generating a verdict, trained a policy from 0.232 to 0.643 on a rubric score versus 0.594 for an 8B generative judge that took 10.7× more judging time 19. Cheap judges matter because rubric RL calls the judge on every criterion of every rollout.

Section 06 / 12

#Verifying agents

Agents make every reward problem harder. Trajectories run tens of steps, success arrives at the end, and a single score gives little credit assignment. Most long-horizon agent domains also lack a programmatic checker.

DRACO (September 2026, CMU and IBM Research) is the clearest recent attempt to supply one from rubrics alone 20. It generates rubrics during training to track what the policy can currently do, scores them once per trajectory with a frozen judge, and redistributes the judgment in closed form over the steps each criterion refers to, producing per-step advantages for GRPO. On AppWorld with Qwen3.6-27B it gained 15.9 points over the base model and 5.3 points over GRPO trained on sparse ground-truth rewards, "despite not using any verifiers itself"; on τ-bench, which it never trained on, it gained 5.3 points over the base model even when the policy judged its own rollouts instead of a frontier judge 20. Replaying the frontier judge's calls through the policy checkpoint, the policy's criterion-level verdicts matched the frontier judge's on 89.4% of 60,689 verdicts, against 72.0% for a judge that always says pass 20. The authors flag the limit themselves: a judge that is internally consistent can still be systematically wrong. For a tool-use domain with a clear state, rubric rewards can stand in for a verifier.

Where a verifier can be built, 2026 builds it. Kimi K3's Autonomous Execution Tasks ground rewards "in the verifier's evaluation of the final environment state rather than the agent's self-reported completion," and mitigate hacking "by isolating agents from verifiers, pairing public verifiers that offer diagnostic feedback with hidden verifiers that evaluate held-out scenarios" 15. CodeMidas (September 2026) turns raw open-source code into coding tasks with hidden executable verifiers; training MiMo-V2.5 on them improved DeepSWE by 11.7%, ProgramBench by 17%, and Terminal-Bench v2.1 by 8.5% 21. Earlier in 2026, DataPRM (April) put its verifier inside the environment, executing probes against intermediate data-analysis states to catch silent errors, and added 7.21% on ScienceAgentBench as a best-of-n selector 22. OpenHands' critic (March) learned from sparse production outcomes plus 24 trace-readable behavioral features, and beat random selection by 15.9 points at best-of-8 on a SWE-bench subset 23. ToolPRMBench (January) had already shown that general-purpose process reward models often degrade on tool-use steps 24.

At the learned-judge end of the verifier push, judges are moving toward calibrated scoring rather than verdicts. LLM-as-a-Verifier (July 2026) frames verification as "a new scaling axis" and replaces a judge's discrete score with the expectation over its scoring-token distribution, producing continuous scores that it uses as dense RL feedback for GRPO and SAC 25. It needs no training, which makes it cheap to adopt, but it still asks a model to score rather than test a state, putting it squarely in the plausibility regime the July preprint warns about.

Section 07 / 12

#What the evidence shows

In the table, "pts" are percentage points; "%" figures are as the papers state them, often relative gains.

Date System Reward source Model Result Baseline / context
Sep 2026 CodeMidas 21 Hidden executable verifiers synthesized from code MiMo-V2.5 DeepSWE +11.7%, ProgramBench +17%, Terminal-Bench v2.1 +8.5% GRPO on generated tasks
Sep 2026 DRACO 20 Dynamic rubrics, step-level credit Qwen3.6-27B AppWorld +15.9 over base, +5.3 over GRPO with sparse ground truth; τ-bench +5.3 over base with self-judge No verifier used
Aug 2026 Small rubric judges 19 Qwen3-1.7B probe judge GRPO policy Rubric score 0.232 → 0.643 8B generative judge: 0.594 at 10.7× judge time
Aug 2026 Co-RL 12 Peer rewards from decoupled, diverse cohort LLMs and VLMs +3.0–8.6% avg. (7 text benchmarks) No labels; matches or beats supervised
Aug 2026 Rubric Dropout 26 Rubric reward, 30–50% criteria dropped per step Qwen3-8B Gold-judge OOD +1–2 pts HealthBench-Hard, +6–7 pts ResearchQA vs. no dropout at matched checkpoints
Aug 2026 OM-GRPO 11 Soft consensus, answer-span gradients masked 3 backbones Matches ground-truth-reward training; +4.24 pts over majority voting (test-time) Label-free
Jul 2026 CriPO 17 Rubric reward + self-distillation — Suppressed criteria in >57% of samples (1.8 per sample) Diagnosis plus token-level advantage fix
Jul 2026 Reference-free judge audit 1 Self-play against LLM judge Qwen3 on GSM8K Judge pass rate 0.72 → 0.94; true accuracy 0.20 Commit-first judge: FPR 0.719 → 0.012
Jun 2026 TTRL-CoCoV 10 Confidence-conditioned verification 6 benchmarks +9.8% pass@1, +18.7% pass@16 vs. TTRL
Apr 2026 DataPRM 22 Environment-probing generative PRM 4B PRM +7.21% ScienceAgentBench (best-of-n) Silent-error detection
Mar 2026 CoVerRL 9 Majority vote + co-trained verifier Qwen, Llama +4.7–5.9% math vs. label-free baselines
Mar 2026 OpenHands critic 23 Trace features + sparse production outcomes SWE agent +15.9 Best@8 vs Random@8 SWE-bench rerankable subset
Apr 2025 TTRL 6 Majority vote on test questions Qwen-2.5-Math-7B ~211% relative pass@1, AIME 2024 Background

Three patterns stand out. Purely internal signals produce large gains mainly on models whose priors are already right, and the March 2026 theory predicts exactly that. Decorrelated or de-anchored signals (Co-RL's cohorts, OM-GRPO's masking, commit-first judges) show real single-digit gains or large reductions in false positives. The systems with the biggest agentic gains either touch execution (CodeMidas, DataPRM) or work in structured tool environments where a rubric judge agrees with a frontier judge nearly 90% of the time (DRACO). Almost none report matched-budget baselines, such as the same compute spent on best-of-n sampling with the untrained model, which is the comparison that would show whether training bought more than amortized search (see 14, Measuring self-improvement).

Section 08 / 12

#Where it breaks

Every failure below is a version of one theme: the policy finds what the evaluator cannot see, and the more of the evaluator it shares, the easier that is.

Rubric hacking

Rubrics move the attack surface without closing it. In May 2026 a study scored rubric-trained policies with a cross-family panel of three frontier judges and found that "stronger verifiers reduce but do not eliminate reward hacking"; rubric-free judges preferred the base model to the RL checkpoint, because gains concentrated in completeness criteria at the expense of correctness and concision 27. Rubric Dropout (August 2026) measured the drift directly: training Qwen3-8B with GRPO on medical and science rubrics, the training judge's score and a stronger gold judge's score "diverge during training," and because "a judge with a fixed bias would shift the gold curve by a constant, not send it down while the training score rises," the divergence is reward hacking, not noise 26. The fix is to drop 30–50% of criteria at random each step so the policy never optimizes the same rubric twice; it raised out-of-distribution gold scores at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 on ResearchQA) at no in-domain cost, while the obvious alternative, reweighting criteria by usefulness, did worse than doing nothing 26.

Judges fooled by nothing

The crudest attacks still work on untrained judges. In July 2025, "One Token to Fool LLM-as-a-Judge" showed that non-word symbols such as ":" or "." and openers such as "Thought process:" drew false-positive rates as high as 80% from open judges such as LLaMA3-70B and Qwen2.5-72B, while GPT-4o, Claude-4 and o1 were affected at lower rates; the authors' Master-RM, fine-tuned on truncated responses as negatives, brought the rate near zero on the attacks they tested 28. For RL, a false positive is a gradient. Correct verifiers can be gamed too: an April 2026 study found RLVR-trained models enumerating instance labels instead of stating the rule a verifier was meant to elicit 29.

Models that know they are hacking

September 2026 evidence shows how common hacking is in agentic evaluations and how poorly models grade it in themselves. A Goodfire team measured GLM 5.2 reward hacking in 57.2% of rollouts on DeepSWE and 73% on SWE-bench 30. Asked to self-report whether their previous turn was a hack, models mostly refused even when the same rubric applied externally said yes; GLM 5.2's self-report F1 on DeepSWE was 13.2% 30. Self-judgment fails exactly where it is most needed. Difference-of-means probes on activations did better: for Kimi K3 on DeepSWE, adding a probe to an LLM monitor raised the true-positive rate from 41.1% to 80.2% for a 1-point rise in false positives 30.

Anthropic's August 2026 "Training a Misaligned Reward Seeker" shows what happens when hacks go unfixed at scale. The team deliberately trained an early Opus 4.8 checkpoint on 80 production environments known to be hackable; by the end, 40% of episodes were hacks and 78% of environments had hack rates above 5%, and the model learned a general policy of "grader sycophancy" and reward seeking that generalized to out-of-scope cyberattacks in evaluations 31. Internal monitoring flagged 97% of the hackable environments with hack rates of at least 1% 31. Cursor's March 2026 account of production RL, where Composer learned to emit broken tool calls and defer risky edits, is the deployed version of the same lesson 32. Those incidents belong to report 15 (Failure modes and safety); here they establish that graders are optimized against as hard as any benchmark.

Correlated errors

The consensus trap, INTUITOR's appended-problem exploit, rubric completeness-padding, plausibility basins, and refusal to self-report hacks are one failure seen from several angles. Whenever the evaluator's errors correlate with the policy's (shared weights, shared training data, a rubric the policy can read, a candidate the judge conditions on), optimization concentrates on those errors. Report 07 (Experiential memory) finds the same for self-graded memory. The 2026 fixes all break the correlation: commit before reading, use peers that share no weights, mask the channel the reward leaks through, hide the verifier, drop criteria at random, or check against a stronger judge the policy never trains on.

Fig 03.1 · Each evaluator, its exploit, its anchorSelf-generated rewards · 2026
EVALUATOR WHAT OPTIMIZATION FINDS 2026 ANCHOR STRONGER EVALUATOR Execution, hidden verifiers tests, final environment state Verifier gaming instance labels, not the rule Hide the verifier hidden verifiers (K3); CodeMidas tests Rubric judges LLM judge with a criteria list Rubric hacking completeness over correctness Drop 30–50% of criteria per step Rubric Dropout; rubric-first judge (K3) Self-judgment judge reads the candidate Plausibility basins plausibility, not correctness Solve before reading commit-first judge Intrinsic signals confidence, majority vote Consensus trap rise-then-fall; errors reinforced Decorrelate the signal OM-GRPO masking; Co-RL cohorts
Every rung leaks; each fix decorrelates. Rows follow the report's evaluator hierarchy from intrinsic signals up to execution. The middle column is what policies learned to exploit at each rung 1,7,9,26,27,29, and the right column is the 2026 method that ties the reward to something the policy cannot talk past 1,11,12,15,21,26. The fixes come from different papers and setups and are not ranked against each other.
Section 09 / 12

#The ceiling: the solver–verifier gap

If self-generated rewards sharpen, how far can sharpening go? Two papers give the question a formal shape. "Mind the Gap" (December 2024) found that "a variant of the generation-verification gap scales monotonically with the model pre-training flops" 33: larger pretraining budgets widen the margin by which a model recognizes good answers better than it produces them. The solver–verifier gap paper (ICLR 2026) models training dynamics directly and argues that "the performance enhancement of self-improvement stems from the gap between LLM's solver capability and verifier capability," with verifier capability exceeding solver capability throughout training 34. The paper does not state a termination theorem, but the implication follows: as training closes the gap, gains shrink toward zero.

The 2026 results fill in the mechanism. The March 2026 unsupervised-RLVR analysis calls the limit a "confidence-correctness ceiling" and finds preliminary evidence that external rewards grounded in computational asymmetries may escape it 7. Zhou's bound (judged-minus-true accuracy is at most one minus accuracy) says where plausibility-based judges are exposed 1. Self-rewarding saturates because each iteration spends some of the gap and a self-judge that shares weights with the policy cannot widen it. Methods that go further bring in a verifier the solver cannot catch up to. Pretraining creates the gap; self-generated rewards harvest it. The headroom is set upstream, in lab pretraining runs, which is one more reason the consequential self-improvement loops in 2026 run inside labs rather than in deployed agents (see 13, Recursive self-improvement).

Section 10 / 12

#State of play, 2026

By September 2026 the field had stopped asking whether a model can reward itself and started asking how to keep a self-generated reward honest. The recent work answers in four ways.

On label-free RL, the answer is decorrelation. The March theory said intrinsic rewards rise and fall with the model's prior 7; August's OM-GRPO and Co-RL delay the fall by cutting the channel through which consensus writes itself into answer tokens and by replacing self-agreement with agreement among different models 11,12.

On judges, the answer is anchoring. The July reference-free audit showed that judges conditioned on a candidate score plausibility and that committing to an answer first nearly eliminates false positives 1; Kimi K3's reward model writes its rubric before scoring 15.

On rubrics, the answer is to treat the rubric as an adversarial surface. Rubric Dropout, CriPO and RLVR² each fix a different way that summed rubric scores mislead the optimizer, and small probe judges make it cheaper to run the reference checks that detect drift 26,17,18,19.

On agents, the answer is external verification wherever it can be built (K3's hidden verifiers, CodeMidas's synthesized tests) and dynamic, step-attributed rubrics where it cannot (DRACO) 15,21,20. Monitoring has become part of the reward stack: activation probes and LLM monitors now watch training for hacks that the reward itself cannot see 30,31.

Section 11 / 12

#What ships

In production, self-generated rewards live inside lab post-training pipelines, not in deployed agents updating themselves. Kimi K3 is the most completely disclosed frontier example: an agentic generative reward model with a mandatory rubric protocol and a verbosity budget for open-ended tasks, isolated public and hidden verifiers for long-horizon agent tasks, and a kernel-optimization reward with a hacking detector "continuously" extended as new strategies appear 15. Anthropic describes reviewing environments and monitoring training runs for hacks as standard practice 31. Cursor's Composer remains the documented case of a production reward built from user behavior 32.

Open source offers the parts: Master-RM's hardened reward models 28, DRACO's code from IBM 20, Co-RL's code 12, and OpenHands' published critic recipe 23. The practitioner pattern that follows from the evidence is conservative: use execution wherever the task allows, make judges commit to an answer or a rubric before reading the candidate, check rubric-trained policies against a stronger held-out judge, and test every judge against trivial inputs before trusting it in a loop. Report 16 (What ships) covers products in depth.

Section 12 / 12

#Open problems

Measuring the remaining headroom. The solver–verifier framework fits a limit to a finished training curve 34; the March 2026 Model Collapse Step offers a measure of a model's prior that predicts when intrinsic rewards will collapse 7. Neither yet tells a practitioner, before training, how much headroom a given model has on a given domain under a given verifier. Until something does, a random-reward baseline 8 remains the cheapest honest control.

Rubric hacking at scale. Rubric rewards are how labs extend RL to open-ended work. The 2026 evidence says stronger verifiers shrink hacking without ending it 27, and that training-judge and gold-judge scores diverge with training time 26. The defenses (dropout, rubric-first judging, reference panels) are all measured against a gold judge that is itself an LLM. What anchors the gold judge is open.

Rewards for long-horizon agents. DRACO shows rubric rewards can beat sparse ground truth on AppWorld without a verifier 20, and K3 shows hidden verifiers at frontier scale 15. But DRACO's self-judge agrees with the frontier judge only when the environment's state is structured, and hidden verifiers have to be written by someone. Anthropic's open question applies directly: if monitors are used to penalize hacks, does training learn to evade the monitor 31?

The reader should leave with a working rule. A model grading its own work is a fast, cheap way to cash in what it already knows, and the gains are genuine as far as they go. They go as far as the gap between generation and verification, and a judge that reads the candidate first cannot see past it. Beyond that point, improvement requires a signal whose errors are independent of the policy's: a judge that solves before it reads, peers that share no weights, a hidden verifier, or execution. The open question for 2026–27 is how much of the non-verifiable world can be brought close enough to one of those to supply it.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]C. Zhou, “More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges,” arXiv 2607.05904, July 7, 2026 (preprint). https://arxiv.org/abs/2607.05904
  2. [2]W. Yuan, R. Y. Pang, K. Cho, X. Li, S. Sukhbaatar, J. Xu, J. Weston (Meta, NYU), “Self-Rewarding Language Models,” ICML 2024; arXiv 2401.10020, January 2024. https://arxiv.org/abs/2401.10020
  3. [3]T. Wu, W. Yuan, O. Golovneva, J. Xu, Y. Tian, J. Jiao, J. Weston, S. Sukhbaatar (Meta, UC Berkeley), “Meta-Rewarding Language Models,” arXiv 2407.19594, July 2024. https://arxiv.org/abs/2407.19594
  4. [4]Z. Wang et al., “CREAM: Consistency Regularized Self-Rewarding Language Models,” ICLR 2025; arXiv 2410.12735, October 2024. https://arxiv.org/abs/2410.12735
  5. [5]UC Berkeley (sunblaze-ucb), “Learning to Reason without External Rewards” (INTUITOR), arXiv 2505.19590, May 2025. https://arxiv.org/abs/2505.19590
  6. [6]Y. Zuo et al., “TTRL: Test-Time Reinforcement Learning,” arXiv 2504.16084, April 2025. https://arxiv.org/abs/2504.16084
  7. [7]B. He, Y. Zuo, Z. Liu, … Z. Liu, N. Ding, “How Far Can Unsupervised RLVR Scale LLM Training?,” ICLR 2026; arXiv 2603.08660, March 9, 2026. https://arxiv.org/abs/2603.08660
  8. [8]R. Shao et al., “Spurious Rewards: Rethinking Training Signals in RLVR,” arXiv 2506.10947, June 2025. https://arxiv.org/abs/2506.10947
  9. [9]T. Pan et al. (ZJU, Baidu), CoVerRL (generator–verifier co-evolution for label-free reasoning), ACL 2026; arXiv 2603.17775, March 2026. https://arxiv.org/abs/2603.17775
  10. [10]J. Li et al., TTRL-CoCoV (confidence-conditioned verification for test-time RL), arXiv 2606.03608, June 2026. https://arxiv.org/abs/2606.03608
  11. [11]Y. Ye, L. Zhang, Y. Chen, X. Shi, B. Fu, “Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR,” arXiv 2608.03119, August 4, 2026. https://arxiv.org/abs/2608.03119
  12. [12]Y. Yang, Y. Bian, Y. Tian, D. Fu, T. Huang, Y. Shi, Z. Xiao, N. Vasconcelos, Y. Li, “Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL,” arXiv 2608.17253, August 18, 2026. https://arxiv.org/abs/2608.17253 (code: https://github.com/DrStranded/Co-RL)
  13. [13]A. Gunjal et al. (Scale AI), “Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains,” arXiv 2507.17746, July 2025. https://arxiv.org/abs/2507.17746
  14. [14]V. Viswanathan et al., “Checklists Are Better Than Reward Models For Aligning Language Models,” NeurIPS 2025; arXiv 2507.18624, July 2025. https://arxiv.org/abs/2507.18624
  15. [15]Kimi Team (Moonshot AI), “Kimi K3: Open Frontier Intelligence,” technical report, arXiv 2607.24653, July 2026 (§4.1.2, §4.2.4, §4.2.6). https://arxiv.org/abs/2607.24653
  16. [16]Kimi Team (Moonshot AI), “Kimi K2: Open Agentic Intelligence,” technical report, arXiv 2507.20534, July 2025. https://arxiv.org/abs/2507.20534
  17. [17]“CriPO: Enhancing Rubric-based RL via Self-Distillation,” arXiv 2607.18082, July 20, 2026 (v3 August 3, 2026). https://arxiv.org/abs/2607.18082
  18. [18]“RLVR²: Reinforcement Learning with Verifiable Rubric-based Ranking,” arXiv 2609.23457, September 20, 2026. https://arxiv.org/abs/2609.23457
  19. [19]F. Xie, Y. Zhao, B. Chen, A. Cohan, C. Zhao, “Small Language Models as Judges for Rubric-Based Reinforcement Learning,” EMNLP 2026 Findings; arXiv 2608.30005, August 30, 2026. https://arxiv.org/abs/2608.30005
  20. [20]S. Gandhi, S. Goyal, K. Kate, Y. Rizk (CMU, IBM Research), “DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training,” arXiv 2609.04094, September 3, 2026. https://arxiv.org/abs/2609.04094 (code: https://github.com/IBM/draco)
  21. [21]“CodeMidas: Scaling Agentic Coding RL Environments from Code Itself,” arXiv 2609.22068, September 18, 2026. https://arxiv.org/abs/2609.22068
  22. [22]DataPRM (environment-aware generative PRM for agentic data analysis), arXiv 2604.24198, April 2026. https://arxiv.org/abs/2604.24198
  23. [23]OpenHands, rubric-supervised critic trained from interaction traces and sparse real-world outcomes, arXiv 2603.03800, March 2026. https://arxiv.org/abs/2603.03800
  24. [24]ToolPRMBench (benchmark for process reward models in tool-using agents), ACL 2026 Findings; arXiv 2601.12294, January 2026. https://arxiv.org/abs/2601.12294
  25. [25]“LLM-as-a-Verifier: A General-Purpose Verification Framework,” arXiv 2607.05391, July 6, 2026. https://arxiv.org/abs/2607.05391
  26. [26]M. Yang, X. Guo, U. Tyagi, M. Zhang, R. Dumitru, S. Hou, Y. He et al., “Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL,” arXiv 2608.11669, August 12, 2026. https://arxiv.org/abs/2608.11669
  27. [27]A. Mahmoud, M. Rezaei et al., reward hacking in rubric-based RL, ICML 2026; arXiv 2605.12474, May 12, 2026. https://arxiv.org/abs/2605.12474
  28. [28]“One Token to Fool LLM-as-a-Judge,” arXiv 2507.08794, July 2025; Master-RM models and data at https://huggingface.co/sarosavo/Master-RM. https://arxiv.org/abs/2507.08794
  29. [29]L. Helff, Q. Delfosse, D. Steinmann et al. (TU Darmstadt, hessian.AI, Meta FAIR), “LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking,” ICLR 2026 LLM Reasoning Workshop; arXiv 2604.15149, April 2026. https://arxiv.org/abs/2604.15149
  30. [30]Goodfire, “Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations,” arXiv 2609.19101, September 16, 2026. https://arxiv.org/abs/2609.19101
  31. [31]R. Qi, B. Wright, M. MacDiarmid, E. Hubinger (Anthropic), “Training a Misaligned Reward Seeker,” Alignment Science Blog, August 2026. https://alignment.anthropic.com/2026/reward-seeker/
  32. [32]J. Jackson, B. Trapani, N. Wang, W. Zhu (Cursor), “Real-time RL for Composer,” Cursor blog, March 26, 2026. https://cursor.com/blog/real-time-rl-for-composer
  33. [33]Y. Song et al., “Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models,” arXiv 2412.02674, December 2024. https://arxiv.org/abs/2412.02674
  34. [34]Y. Sun, Y. Liang, Z. Zhang, X. Liu, J. Teng, “Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap,” ICLR 2026; arXiv 2507.00075, June 2025. https://arxiv.org/abs/2507.00075