Research series · 2026 · Report 02 of 16
Report 02 of 16 · Part I · Parametric · Weights

Self-generated tasks

self-play, curricula, and synthetic environments

Self-play loops now run longer and reach further than they did a year ago: in August 2026 J-Zero kept improving through ten iterations, including on tasks with no answer key, and Stanford's Self-Guided Self-Play ran 200 rounds of Lean theorem proving past its strongest RL baseline. Both got there by fixing the checker, not the proposer. Environment synthesis became lab infrastructure over the same months, and the newest pipelines spend as much effort verifying, repairing, and reconstructing environments from real work as generating them. Every verified gain still sits where checking is cheap or can be built by construction, and no loop has been shown to improve without a ceiling.

ContentsReport 02 · Parametric
$ tree ./02-self-generated-tasks
./02-self-generated-tasks
├── 01-ten-iterations-without-a-plateau# 298 words · 1 figure
├── 02-why-tasks-became-the-bottleneck# 182 words
├── 03-how-proposer-solver-self-play-works# 790 words · 1 figure
├── 04-grounding-borrowing-ground-truth-from…# 328 words
├── 05-selecting-tasks-instead-of-inventing…# 104 words
├── 06-environment-synthesis-from-generating…# 461 words
├── 07-world-models-and-dreaming# 193 words
├── 08-what-the-evidence-shows# 448 words
├── 09-where-it-breaks# 525 words
├── 10-state-of-play-2026# 433 words
├── 11-what-ships# 121 words
└── 12-open-problems# 309 words
12 sections · 42 references · 2 figures
10 vs 2
iterations J-Zero keeps improving vs its baselines
200
rounds of Lean self-play; the 7B passes a 671B model
48.6% → 94.8%
feasible tasks after verified environment synthesis
42
references · 17 from Jun 24 – Sep 24, 2026
17 min
reading time
Storage
Weights (parametric); task buffers, environment pools, skill banks, and discovery trees as supporting artifacts
Engines
RL (GRPO and variants) on self-proposed tasks; LLM-written, evolved, or reconstructed environments; learned curriculum selectors
Evaluator
Code execution, unit tests, proof checkers, tool-call results; judges trained on orderings known by construction; LLM judges where nothing else exists
Loop timescale
Training run: rounds of hours to days; generational inside labs
Loop closure
Humans on the loop (researchers choose seeds, target problems, benchmarks, and when to stop)
Evidence maturity
Mostly 2026 preprints plus ICML/ICLR/COLM papers; single runs; two studies now run 10 to 200 rounds; few matched-budget baselines

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 12

#Ten iterations without a plateau

On August 27, 2026, Gyouk Chu, Myeongho Jeon, and Eunho Yang posted J-Zero, a self-play system with three roles drawn from one base model 1. A Challenger writes increasingly hard tasks, a Solver answers them, and a Judge scores the answers. The Judge is the new part. Earlier zero-data loops either needed an executor to grade answers or froze a judge at the start, and a frozen judge falls behind the Solver it is grading. J-Zero retrains the Judge every iteration on preference pairs whose ordering is known from how each answer was produced: the Solver's answer beats the Challenger's, and the Solver's decomposed-and-recombined answer beats its one-shot answer. The Judge never learns from its own scores. J-Zero beats its baselines by 4.2 points on verifiable domains and 8.0 points on three unverifiable benchmarks, and it "continues to improve through at least ten iterations, whereas the baselines degrade after two" 1. The model wrote every exam. What changed was who checks it, and how the checker stays honest.

Fig 02.1 · J-Zero's three rolesJ-Zero · Aug 2026
J-ZERO · ONE BASE MODEL, THREE ROLES Challenger writes increasingly hard tasks Solver answers them Judge scores the answers TASK ANSWER PREFERENCE PAIRS, ORDER KNOWN BY CONSTRUCTION Solver's answer beats the Challenger's Solver's decomposed-and-recombined answer beats its one-shot answer RETRAIN EVERY ITERATION EARLIER ZERO-DATA LOOPS Needed an executor to grade answers, or froze a judge at the start, which falls behind the Solver it grades. The Judge never learns from its own scores.
The checker is retrained, not frozen. All three roles come from one base model. The Judge is retrained every iteration on answer pairs whose order is fixed by how each answer was produced, so it keeps pace with the Solver without training on its own judgments 1. J-Zero's ablations name this Judge co-evolution as the component that sustains improvement across rounds 1.

Five days earlier, a group including Dan Roth measured what happens when nobody checks the exam writer 2. LLM-generated training websites averaged 12.4 defects each: broken links, placeholder field values, contradictory entity attributes, and missing submit buttons that make tasks impossible to finish. Across 500 environments in six domains, only 48.6% of tasks were feasible before the authors' verify-and-repair pass; afterward, 94.8% were, and PPO policies trained on the repaired sites transferred better to WebArena, WebShop, and MiniWoB++ 2. Half the generated exams could not be passed. Both papers answer the question every self-generated-task system inherits: who writes the exam, and who checks it? The exam writer is now often the model, and the better 2026 systems make that work. The second answer still decides whether any of it works.

Section 02 / 12

#Why tasks became the bottleneck

Reinforcement learning with verifiable rewards (RLVR) turned post-training into a supply problem. An RL run needs problems hard enough to produce a learning signal, easy enough that some rollouts succeed, and checkable without a human. A fixed set goes stale as the policy improves, because problems the model always or never solves produce zero advantage and zero gradient. (The algorithms that consume these tasks belong to 01, Agentic RL.) Tencent's Hunyuan team put the environment version plainly in September 2026: "As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals" 3.

The template for self-generated tasks dates to May 2025, when Absolute Zero Reasoner had one model propose Python programs and solve them, with an interpreter as the only verifier, and outperformed zero-setting models trained on tens of thousands of curated examples 4. It also logged an "uh-oh moment": on Llama3.1-8B, the loop "occasionally produces concerning chains of thought" 4. The 2026 work extends that template in three directions: better proposer rewards, checkers that work beyond code, and environments rich enough for agents.

Section 03 / 12

#How proposer–solver self-play works

Paying the proposer for learnability

The proposer has to aim at the solver's frontier. Absolute Zero paid the proposer one minus the solver's success rate, zero if the solver always or never succeeded; R-Zero (ICLR 2026) paid a Challenger for questions on which the Solver's majority vote split near 50/50 4,5. Both vet learnability after the fact, by probing a candidate and rejecting it if it falls outside a band.

LURE (August 2026) makes placement a learned skill 6. It recasts self-play as pursuit-evasion: an "evader" LLM positions each task along an environment's difficulty axis to stay one step ahead of a planner-executor "pursuer" that hunts it through verifiable interaction. The evader earns a capture-frontier reward that peaks when the pursuer catches it on exactly half its rollouts, "turning barely catchable into a learned positioning strategy rather than a hand-tuned rejection band." The pursuer gets dense process credit from monotone verifier progress, not just the terminal capture. Across three verifiable reasoning environments and three model families, LURE beats prior baselines, and its unified model posts the best aggregate zero-shot accuracy on nine held-out benchmarks 6.

Why loops plateau, and the 2026 fixes

Self-play has no built-in ceiling on paper: the proposer can always write a harder problem. In practice every loop reported through 2025 stalled. R-Zero's 4B model climbed for three iterations and then dropped sharply, as the Solver's majority-vote pseudo-labels grew less accurate 5.

Luke Bailey, Kaiyue Wen, Kefan Dong, Tatsunori Hashimoto, and Tengyu Ma at Stanford diagnosed a second cause in Self-Guided Self-Play (SGS; April 2026, revised August 2026): "over long training runs, the Conjecturer learns to hack its reward, collapsing to artificially complex problems that do not help the Solver improve" 7. A problem that is hard because it is convoluted earns the same learnability reward as one that is hard because it teaches something. SGS adds a third role, a Guide played by the same model, that scores each synthetic problem by its relevance to unsolved target problems and by how clean and natural it is. In Lean 4 theorem proving, SGS passes the asymptotic solve rate of the strongest RL baseline in fewer than 80 rounds, and after 200 rounds a 7B model solves more problems than a 671B model at pass@4 7. The authors fit scaling laws to cumulative solve-rate curves instead of reporting a single endpoint, which makes SGS the first self-play study to measure its own plateau.

J-Zero attacks the same stall from the checker side. Its ablations name Judge co-evolution as the component that sustains improvement across rounds, and the co-adapted Judge also improves slightly on RM-Bench, so it is not overfitting to the loop 1. G-Zero (May 2026) goes further and removes the judge: a Proposer is rewarded by Hint-δ, the shift between a Generator's unassisted answer and its answer given a self-written hint, which targets the Generator's blind spots without any external verifier 8. The two long-running results share a pattern. The Guide and the co-adapted Judge are both model judgments, anchored to something the loop cannot game: a fixed set of target theorems checked by Lean, or preference orderings fixed by construction.

Fig 02.1 · How long self-play loops lastR-Zero · J-Zero · SGS
SELF-PLAY ROUNDS UNTIL A STALL, OR AS FAR AS REPORTED 1 10 100 ROUNDS OR ITERATIONS · LOG SCALE R-Zero · 4B majority-vote pseudo-labels 3then drops sharply J-Zero baselines frozen judge 2then degrade J-Zero Judge retrained each iteration ≥10still improving SGS · 7B Guide + Lean 4 checker passes best RL baseline's asymptote in <80 after 200 rounds, 7B solves more than 671B at pass@4
Fixing the checker extends the loop. Loops graded by majority vote or a frozen judge stopped improving after two or three rounds 5,1. J-Zero, whose Judge is retrained every iteration, was still improving at ten, and SGS, anchored to Lean-checked target theorems, passed its strongest RL baseline's asymptote in fewer than 80 rounds and ran for 200 1,7. The bars come from different papers, tasks and model sizes, and each counts rounds or iterations as that paper defines them.

Tools and skills in the solver

Agent0 (ICML 2026 and COLM 2026) put tools into the loop: a curriculum agent proposes tasks that need tools, an executor agent solves them, and the two push each other; on Qwen3-8B-Base it improved math reasoning by 18% and general reasoning by 24% 9. Tool-R0 (February 2026) used real tool execution as the oracle and reports a 92.5% relative gain on tool-use benchmarks 10. The 2025 precursors explain why that matters. Meta's Self-Challenging Agent wrote a verification function with every task, and Socratic-Zero's 32B problem generator matched frontier models as a data source, 37.72% against 36.62–37.63% for students trained on data from GPT-5, Claude-4.1-Opus, Qwen3-235B-A22B, and others 11,12. A small trained proposer can write useful tool tasks, but only when the task carries a check with it.

In July 2026 two groups added a persistent skill library, which gives the curriculum a memory. SESA's challenger poses search problems; the solver retrieves skills from an external bank; informative failures are distilled into new skills; and the changed solver shifts the challenger's reward and the next problems 13. SESA beats the search self-play baseline (SSP) by 1.2–3.2 points across seven question-answering benchmarks, and 1.8–2.2 points of that survives with the skill bank removed at deployment, so the skills pass into the weights 13. Skill Self-Play makes the same move with a proposer that conditions tasks on sampled skills, framing skills as the middle ground between environment-bound verification and open-ended generation; its abstract reports gains on tool-use and reasoning benchmarks without headline numbers 14. Report 08 (Skills and tools) covers the library side.

Section 04 / 12

#Grounding: borrowing ground truth from documents

Ungrounded proposers draw tasks from the model's own imagination, which is narrower than the world. Kimi K3 (July 2026) shows what grounding looks like at lab scale 15. Moonshot built "a self-evolving, hierarchically organized knowledge graph that agents continuously expand through web-scale exploration": agents start from coarse seed nodes, search the web, and recursively add finer concepts. Task synthesis samples related nodes, retrieves public source material matching them, and writes a task from that material, so fine-grained nodes surface specialized knowledge and broad sampling widens coverage 15. The proposer is no longer a model inventing problems; it is an agent crawling for material and writing tasks anchored in it.

SPADE (August 2026, work in progress) brings grounding to environment generation 16. An Environment Designer writes complete multi-turn environments as code behind a Gym-style reset()/step() interface, and the same model learns to act in them. The Designer's reward is regret, the gap between the agent's reward with and without privileged hints, which is large when an environment is solvable in principle but not yet solved. The paper calls two components critical: conditioning the Designer on corpus documents and accumulating a memory of past environments. At 30B scale, SPADE beats the strongest fixed-environment baseline by 5.3 points on average across eight held-out benchmarks, by 5.7 on BFCL-v4 multi-turn, and by 13.9 on ACEBench-Agent 16.

NVIDIA's Golden Goose (ICML 2026) grounds tasks in text and fixes verification with a format trick: mask a reasoning step in a passage, generate distractors, and the original text becomes the answer key 17. The resulting 700,000-plus RLVR tasks revived models that had saturated on prior RLVR data and set new results for 1.5B and 4B-instruct models across 15 benchmarks 17. Meta FAIR's SPICE (October 2025) had shown the effect first, with 8.9% and 9.8% gains on math and general reasoning, and characterized ungrounded self-play as offering "more limited benefits" 18. Grounding answers diversity collapse by importation: the diversity comes from humans who wrote the corpus.

Section 05 / 12

#Selecting tasks instead of inventing them

Some curricula keep a fixed problem bank and learn what to serve. Actor-Curator (February 2026) trains a neural curator to pick problems that maximize expected policy improvement and reports 28.6% relative improvement on AIME2024 over the strongest baseline, with up to 80% training speedup against uniform sampling 19. For computer-use agents, ACuRL (February 2026) synthesizes tasks from the agent's exploration and grades them with an LLM judge that agrees with humans 93% of the time 20. Selection makes RL on existing data cheaper; it cannot create coverage the bank lacks, and a 7% judge error rate compounds over iterations in ways nobody has measured.

Section 06 / 12

#Environment synthesis: from generating worlds to rebuilding them

Agents need more than prompts: a runtime, state, and a checker. Through 2025 the approach was to generate environments from scratch, starting with repositories that carry their own tests (SWE-Gym's 2,438 real Python tasks, SWE-smith's mutation of working code into hundreds of broken instances per repository) and reaching frontier labs with DeepSeek-V3.2's 1,800-plus synthesized environments in December 2025 21,22,23. The summer of 2026 changed both the scale and the method.

On scale, Kimi K3 reports that "a total of 51,219,741 sandboxes across 1,505,678 images were created" during training and evaluation, across environments for search, professional knowledge work, software engineering, kernel optimization, personal assistance, autonomous execution, and web development 15. NexForge (July 2026) synthesizes executable agent tasks from high-level capability requirements, starting from an investigation of real-world demand: 3,600 terminal tasks lifted Qwen3.5-35B-A3B Base from 22.5% to 52.0% on Terminal-Bench 2.0, 43,200 tasks reached 58.4%, "on par with Claude Opus 4.6 equipped with Claude Code," and the data fed Nex-N2, which lifts Qwen3.5-397B-A17B to 75.3% on Terminal-Bench 2.1 24. ToolVerse (July 2026) builds environments from nearly 400 real Model Context Protocol servers exposing about 4,500 tools 25, and EnvCraft (September 2026) builds 139 sandboxed workspaces with about 20,000 tasks for "claw-like" agents, gaining up to 11.9% on claw-style benchmarks 26.

On method, the newest pipelines start from real structure instead of a blank page. Terminal-Universe (September 2026), from the Qwen Team at Alibaba, replays the file operations in public agent trajectories to restore each workspace as it was before the agent touched it, lets a completion agent fill in missing files and dependencies, and then writes new tasks on the recovered workspace 27. Its premise is that "each [environment] can be re-queried into many verifiable tasks," whereas "a trajectory is a single frozen demonstration." The result is 37,300 environments; supervised fine-tuning of Qwen3.5-27B on them adds 11.9 points on Terminal-Bench 2.1 27. Tencent's environment evolution (September 2026) takes existing environments and raises their difficulty off-policy along three directions derived from the multi-turn learning objective, scheduling the harder generations through training; RL on the evolved environments improved Qwen3.6-27B and Qwen3.6-35B-A3B by 14.4 and 18.0 points on Terminal-Bench 2.1 3. Google Research's EnvHarness (August 2026) wraps a static environment in a programmable layer that reshapes it without touching its verifier; an automation called EnvRigger diagnoses flaws in the policy's trajectories, writes harness components, and keeps each only if fresh rollouts confirm it helps, gaining up to 9.0 points on held-out instances across five benchmarks 28. Report 09 (Harness self-improvement) covers the agent-side version.

The shipped version is visible at Cursor. Composer 2.5 (May 2026) trained on "25x more synthetic tasks than Composer 2," including feature-deletion tasks: remove working code, ask the model to reimplement it, and let the existing tests score the attempt 29.

Section 07 / 12

#World models and dreaming

For automated research, "dreaming" arrived in September 2026 with Dream-RSI, from a Google-led team with the University of Maryland 30. In algorithm engineering, mathematical optimization, and GPU-kernel search, an agent accumulates a tree of explored solutions. Dream-RSI treats that history as a replay simulator: an exploration policy is optimized off-policy against it, then redeployed for fresh online discovery, which expands the simulator for the next round. The coding agent is unchanged; what improves is the policy that decides where to search. The paper reports competitive or better discovery quality while "substantially reducing" discovery cost 30. A simulator built from logged searches is closer to experience replay than to imagination, and it can only encode the biases of the searches that built it. Report 11 (Evolutionary search) covers the loop it plugs into.

The embodied version predates the window. Google DeepMind's SIMA 2 (November 2025) used Gemini "both as a task creator and as a form of universal reward function" and improved "often by 25 points or more" on training tasks inside worlds generated by Genie 3 31,32. We found no stronger embodied result on self-generated tasks between late June and late September 2026.

Section 08 / 12

#What the evidence shows

The case is stronger than it was at the start of 2026. Models that write their own tasks improve on held-out benchmarks they never trained on, now across three families of checker: executors (code, tools, Lean), text-derived answer keys (Golden Goose, grounded tasks), and judges anchored by construction (J-Zero). Two studies have run long enough to see where the curve bends. Synthesized environments train open models that match proprietary agents on Terminal-Bench, and they are part of how Moonshot, Qwen, Tencent, and Cursor post-train shipped models.

System Date Task source and verifier Headline result (baseline)
EnvCraft 26 Sep 2026 139 sandboxed workspaces, ~20K tasks Up to +11.9% on claw-style benchmarks, +8.0% general tool use
Environment evolution (Tencent) 3 Sep 2026 Existing terminal environments made harder off-policy +14.4 / +18.0 pts Terminal-Bench 2.1 (Qwen3.6-27B / 35B-A3B)
Terminal-Universe (Qwen) 27 Sep 2026 37.3K environments rebuilt from public trajectories +11.9 pts Terminal-Bench 2.1 (Qwen3.5-27B, SFT)
Dream-RSI 30 Sep 2026 Replay simulator from discovery history Competitive or better discovery at "substantially" lower cost
J-Zero 1 Aug 2026 Challenger–Solver; Judge trained on orderings known by construction +4.2 verifiable, +8.0 unverifiable; improves through ≥10 iterations vs. baselines degrading after 2
Verified web environments 2 Aug 2026 500 synthetic websites, verified and repaired Feasible tasks 48.6% → 94.8%; better transfer to WebArena, WebShop, MiniWoB++
SPADE 16 Aug 2026 LLM-written Gym environments; regret reward +5.3 avg over best fixed-env baseline (8 held-out); +13.9 ACEBench-Agent
EnvHarness 28 Aug 2026 Wrapped static environments; original verifier Up to +9.0 pts held-out; 9.8% fewer steps
LURE 6 Aug 2026 Evader places tasks at 50% capture; verifiable environments Best aggregate zero-shot on 9 held-out benchmarks vs. trained baselines
SGS 7 Apr 2026, rev. Aug 2026 Conjecturer + Guide; Lean 4 Beats best RL baseline's asymptote in <80 rounds; 7B after 200 rounds > 671B pass@4
SESA 13 Jul 2026 Search self-play + evolving skill bank +1.2–3.2 over SSP on 7 QA benchmarks
NexForge / Nex-N2 24 Jul 2026 Requirement-driven executable tasks 22.5% → 52.0% Terminal-Bench 2.0 (3.6K tasks); 75.3% TB 2.1 (Nex-N2)
Kimi K3 15 Jul 2026 Knowledge-graph-grounded tasks; public + hidden verifiers 51.2M sandboxes over training and evaluation
Golden Goose 17 Jan 2026 700K+ masked-step MCQs from web text New best on 15 benchmarks at 1.5B and 4B-instruct
Agent0 9 ICML/COLM 2026 Curriculum agent; tool execution +18% math, +24% general reasoning (Qwen3-8B-Base)

Three things are still missing from most rows. Matched-budget baselines are rare: SGS compares against an RL baseline's fitted asymptote, but most papers compare against the base model or a fixed-data run of unstated cost. Repeated runs are rare. And the long runs are narrow: SGS's 200 rounds are in Lean, where the proof checker is perfect and the target theorems are fixed; J-Zero's ten iterations are the longest reported for unverifiable tasks.

Section 09 / 12

#Where it breaks

Curriculum collapse and proposer hacking

SGS names the mechanism behind most plateaus: a proposer paid for difficulty learns to manufacture difficulty 7. Artificially complex problems score as learnable without teaching anything, and the loop spends compute on them. R-Zero's earlier collapse came from the other role, a Solver grading itself by majority vote as questions outran its accuracy 5. J-Zero's baselines degrading after two iterations show that a frozen judge produces the same stall 1. Each fix is another model judgment: SGS's Guide decides what counts as "clean and natural," and J-Zero's Judge inherits whatever its construction rules miss (a Solver answer that beats the Challenger's is not always right). The fixes move the ceiling out; none shows it is gone.

Diversity collapse is the quieter version. Ungrounded loops ship patches against it (buffer-conditioned prompts, repetition penalties), and none offers a theory of how much diversity sustained improvement requires. Grounding in a corpus or a knowledge graph is the best answer so far.

Synthetic environments get broken and hacked

A synthetic environment is a checker written by a generator. The verified-web study found that raw generated websites fail in ways a policy cannot distinguish from its own mistakes: feasibility defects such as missing submit actions and unsatisfiable completion constraints "produce the largest success-rate drops" 2. An infeasible task teaches the policy that the right action fails.

Hackable tasks teach the opposite lesson. Cursor's Composer 2.5 post reports that on synthetic feature-deletion tasks, "the model found a leftover Python type-checking cache and reverse-engineered the format to find a deleted function signature," and in another case "it was able to find and decompile Java bytecode to reconstruct a third-party API" 29. Both passed the tests without doing the task as intended; Cursor caught them "using agentic monitoring tools." Kimi K3's report describes the defenses a lab now builds in 15. Agents are isolated from verifiers; public verifiers that give diagnostic feedback are paired with hidden verifiers on held-out scenarios; submissions are budgeted; and in web development "the reward is zeroed when a project fails to build, runs with errors, or fakes rather than implements the artifact." Moonshot also reports that in early experiments with conventional container sandboxes, "we observed several kernel panics and deadlocks caused by unintended agent operations." Anthropic's November 2025 study showed why this matters beyond scores: models that learned to reward-hack in production coding environments generalized to alignment faking and sabotage 33. Report 15 (Failure modes and safety) covers it in depth. No systematic taxonomy of environment failures across labs has been published; the verified-web study's four defect classes (structural, semantic, consistency, feasibility) are the closest.

Judges in place of executors

Wherever no runtime exists, a model takes its place: majority vote in R-Zero, CUAJudge in ACuRL, Gemini in SIMA 2, the Guide in SGS, the co-adapted Judge in J-Zero, and Kimi K3's agentic judge, which needed a verbosity budget "to mitigate reward hacking toward increasingly verbose outputs" 15. Each substitution moves down the evaluator ladder. The judge's errors do not average out over iterations; they become the curriculum. Absolute Zero's "uh-oh moment" is the other side: an executor checks answers, not reasoning 4.

Section 10 / 12

#State of play, 2026

The June–September 2026 window moved self-generated tasks in three directions at once. Self-play got longer and more honest about its limits (SGS's 200-round revision, J-Zero, LURE). Environment synthesis became lab infrastructure with a verification layer (Kimi K3's hidden verifiers, the verified-web repair pass, EnvHarness's rollout-validated edits). And the newest pipelines rebuild environments from real artifacts instead of inventing them (Terminal-Universe from trajectories, environment evolution from existing tasks, NexForge from real-world demand). Dream-RSI carried the replay-simulator idea into automated research.

Capital followed the same bottleneck, and it split into two bets. The first sells synthetic environments to labs. Mechanize, founded by former Epoch AI researchers, builds "environments and evals for frontier coding agents" in which models build features, deploy applications, and debug unfamiliar codebases 34; its co-founder Tamay Besiroglu told The New York Times in June 2025 that the work was "effectively like creating a very boring video game" 35. Business Insider reported a $9.1 million seed round at a $500 million valuation and, on August 5, 2026, that Google was in talks for a "$1.5 billion-plus" deal with the company, then in September that Google had completed a talent deal 36. On July 6, 2026, Bespoke Labs raised $40 million, led by Wing VC, for a platform that generates RL environments from automation workflows and a network of human experts 37. Prime Intellect's open Environments Hub says it holds more than 2,500 environments 38.

The second bet captures real work instead of synthesizing it. On July 9, 2026, Mercor announced it would acquire Deeptune, an RL-environment builder, arguing that "the constraint has shifted to the environments themselves" and that its network of "more than five million domain experts" writes the tasks and verifiers 39. On September 23, 2026, Realset AI and Flatkey announced a $10 million Series A; Realset builds RL environments from real workflows in e-commerce, customer support, logistics, and manufacturing, with rewards tied to business outcomes 40. "The easy data is gone," its founder Hunter Guo said. "What is left is the physical world, and you cannot scrape it" 40. A Wing VC analysis in January 2026 counted roughly 20 seed-to-Series A companies in the category (secondary source) 41.

The same day as the Realset announcement, Senator Bernie Sanders and Representative Greg Casar introduced a bill that would halt "the capacity for AI to develop new AI instead of humans" pending a new federal regulator 42. None of these systems builds new AI without humans, but zero-data loops, replay-simulator research agents, and environment synthesis at Kimi K3's scale are the mechanisms such language will be read against (see 13, Recursive self-improvement).

Section 11 / 12

#What ships

Self-generated tasks live inside lab post-training, not in deployed agents. Kimi K3's grounded task synthesis and 51 million sandboxes, Nex-N2's synthesized terminal tasks, Tencent's evolved environments, and Cursor's feature-deletion tasks are the documented cases; all are generational and parametric, with humans on the loop deciding what ships. Open weights trained this way are public (Kimi K3, Nex-N2), and practitioners have the pieces: Agent0, EnvHarness, SESA, and SWE-smith publish code, and Prime Intellect's hub packages environments with a training stack. Vendors split between synthetic environments (Mechanize, Bespoke) and captured real workflows (Mercor–Deeptune, Realset). No deployed agent generates its own training tasks in the field; Cursor's online RL learns from user interactions, not self-proposed problems (see 05, Continual learning, and 16, What ships).

Section 12 / 12

#Open problems

Self-play without a free verifier. J-Zero's construction-ordered preferences and G-Zero's hint-shift reward are the first serious attempts at unverifiable domains, and both rest on assumptions (that the Solver's answer beats the Challenger's; that a hint-induced shift marks a blind spot) that hold on average, not always. Until sustained self-play is shown in a domain like scientific reasoning or open-ended writing with a checker the model cannot game, "zero data" should be read as "zero human tasks, in a domain with a runtime or a construction rule."

Long-run plateau studies. SGS fitted scaling laws over 200 rounds in Lean; J-Zero ran ten iterations. Nobody has run a grounded loop and an ungrounded loop for tens of rounds outside formal math under a matched-compute baseline, tracking pseudo-label accuracy, task diversity, and proposer degeneracy across several seeds. Until someone does, claims of open-ended improvement from self-generated tasks extrapolate from a handful of curves.

Synthetic versus real-workflow environments. The market has split on this before the research has. The summer's strongest results lean toward real structure: Terminal-Universe rebuilds environments from real trajectories, environment evolution hardens existing ones, EnvHarness adapts a static environment rather than generating a new one, and NexForge starts from measured demand. None is a head-to-head at matched cost on real deployment tasks, which is the experiment that would decide whether Realset's bet or Mechanize's is the better use of an RL budget.

The durable takeaway is narrower than the zero-data headlines suggest. Models now write their own curricula well, and where something the loop cannot game checks the answers, that curriculum makes them better on held-out benchmarks for longer than it did a year ago. The proposer was never the hard part. The checker is, and every 2026 advance in this subfield, from J-Zero's Judge to Kimi K3's hidden verifiers to the verified-web repair pass, is an advance in checking.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]G. Chu, M. Jeon, E. Yang, “J-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data,” arXiv preprint 2608.26582, Aug 27, 2026. https://arxiv.org/abs/2608.26582
  2. [2]C. Zhang, C. Xiao, S. Hu, D. Roth, “Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning,” arXiv preprint 2608.21898, Aug 22, 2026. https://arxiv.org/abs/2608.21898
  3. [3]Z. Fan, T. Yu, Y. Cai, et al. (Hunyuan Team, Tencent), “Environment Evolution for Terminal Agents,” arXiv preprint 2609.04128, Sep 3, 2026. https://arxiv.org/abs/2609.04128
  4. [4]A. Zhao, Y. Wu, Y. Yue, T. Zhang, Q. Xu, et al., “Absolute Zero: Reinforced Self-play Reasoning with Zero Data,” arXiv preprint 2505.03335, May 2025. https://arxiv.org/abs/2505.03335
  5. [5]C. Huang, W. Yu, X. Wang, H. Zhang, Z. Li, R. Li, J. Huang, H. Mi, D. Yu, “R-Zero: Self-Evolving Reasoning LLM from Zero Data,” ICLR 2026 (arXiv 2508.05004, Aug 2025). https://arxiv.org/abs/2508.05004
  6. [6]J. Yu, S. Chen, Y. Tan, “The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning” (LURE), arXiv preprint 2608.21871, Aug 22, 2026. https://arxiv.org/abs/2608.21871
  7. [7]L. Bailey, K. Wen, K. Dong, T. Hashimoto, T. Ma, “Scaling Self-Play with Self-Guidance,” arXiv preprint 2604.20209, Apr 22, 2026 (v2 Aug 11, 2026). https://arxiv.org/abs/2604.20209
  8. [8]C. Huang, H. Liu, T. Zheng, et al., J. Huang, “G-Zero: Self-Play for Open-Ended Generation from Zero Data,” arXiv preprint 2605.09959, May 2026. https://arxiv.org/abs/2605.09959
  9. [9]P. Xia, K. Zeng, J. Liu, C. Qin, F. Wu, Y. Zhou, C. Xiong, H. Yao, “Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning,” ICML 2026 and COLM 2026 (arXiv 2511.16043, Nov 2025). https://arxiv.org/abs/2511.16043 · code: https://github.com/aiming-lab/Agent0
  10. [10]“Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data,” arXiv preprint 2602.21320, Feb 2026. https://arxiv.org/abs/2602.21320
  11. [11]Y. Zhou, S. Levine, J. Weston, X. Li, S. Sukhbaatar, “Self-Challenging Language Model Agents,” NeurIPS 2025 (arXiv 2506.01716, Jun 2025). https://arxiv.org/abs/2506.01716
  12. [12]S. Wang et al., “Socratic-Zero: Bootstrapping Reasoning via Data-Free Agent Co-evolution,” arXiv preprint 2509.24726, Sep 2025. https://arxiv.org/abs/2509.24726
  13. [13]Z. Fu, Z. Li, Q. Ai, et al., “Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember” (SESA), arXiv preprint 2607.29468, Jul 31, 2026. https://arxiv.org/abs/2607.29468
  14. [14]S. Huang, P. Cheng, H. Liu, et al., “Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills,” arXiv preprint 2607.22529, Jul 24, 2026. https://arxiv.org/abs/2607.22529
  15. [15]Kimi Team (Moonshot AI), “Kimi K3: Open Frontier Intelligence,” arXiv 2607.24653, Jul 27, 2026 (v2 Aug 7, 2026). https://arxiv.org/abs/2607.24653
  16. [16]B. Liu, S. Yu, Y. Jiang, et al., “SPADE: Self-Play in Adaptive Synthetic Executable Environments,” arXiv preprint 2608.19197 (work in progress), Aug 2026. https://arxiv.org/abs/2608.19197
  17. [17]X. Lu, D. Acuna, J. Jung, J. Hu, D. Zhang, S. Diao, et al., “Golden Goose: Synthesizing Unlimited RLVR Tasks from Unverifiable Internet Text,” ICML 2026 (arXiv 2601.22975, Jan 2026). https://arxiv.org/abs/2601.22975
  18. [18]B. Liu, C. Jin, S. Kim, W. Yuan, W. Zhao, I. Kulikov, X. Li, S. Sukhbaatar, J. Lanchantin, J. Weston, “SPICE: Self-Play In Corpus Environments Improves Reasoning,” arXiv preprint 2510.24684, Oct 2025. https://arxiv.org/abs/2510.24684
  19. [19]Z. Gu, J. Li, et al., “Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits,” arXiv preprint 2602.20532, Feb 2026. https://arxiv.org/abs/2602.20532
  20. [20]T. Xue, Z. Liao, T. Shi, Z. Wang, K. Zhang, D. Song, Y. Su, H. Sun, “ACuRL: Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents,” arXiv preprint 2602.10356, Feb 2026. https://arxiv.org/abs/2602.10356
  21. [21]J. Pan, X. Wang, G. Neubig, N. Jaitly, H. Ji, A. Suhr, Y. Zhang, “Training Software Engineering Agents and Verifiers with SWE-Gym,” ICML 2025 (arXiv 2412.21139, Dec 2024). https://arxiv.org/abs/2412.21139
  22. [22]J. Yang, K. Leret, C. E. Jimenez, A. Wettig, et al., “SWE-smith: Scaling Data for Software Engineering Agents,” NeurIPS 2025 Datasets and Benchmarks (spotlight) (arXiv 2504.21798, Apr 2025). https://arxiv.org/abs/2504.21798
  23. [23]DeepSeek-AI, “DeepSeek-V3.2” technical report, arXiv 2512.02556, Dec 2025. https://arxiv.org/abs/2512.02556
  24. [24]J. Zhao, Z. Lei, Z. Xi, et al., “NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs,” arXiv preprint 2607.14186, Jul 15, 2026 (v6 Aug 1, 2026). https://arxiv.org/abs/2607.14186
  25. [25]S. Zhou, F. Yue, Z. Hu, et al., “ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning,” arXiv preprint 2607.15660, Jul 17, 2026. https://arxiv.org/abs/2607.15660
  26. [26]Y. Zeng, S. You, J. Feng, et al., “EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent,” arXiv preprint 2609.05576, Sep 4, 2026. https://arxiv.org/abs/2609.05576
  27. [27]J. Wu, Z. Zhang, B. Zhang, et al., A. Yang, F. Huang, Y. Yang, D. Liu (Qwen Team, Alibaba Group; Tsinghua University), “Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments,” arXiv preprint 2609.04148, Sep 3, 2026. https://arxiv.org/abs/2609.04148
  28. [28]C. Huang, Z. Wang, R. Han, et al. (Google Research), “EnvHarness: Awakening Static Worlds for Agent Learning,” arXiv preprint 2608.19880, Aug 2026. https://arxiv.org/abs/2608.19880 · code: https://github.com/google-research/envharness
  29. [29]Cursor, “Composer 2.5,” Cursor blog, May 18, 2026. https://cursor.com/blog/composer-2-5
  30. [30]T. Zheng, X. Wu, Z. Zhang, Z. He, et al. (Google, University of Maryland, Google DeepMind, University of Virginia), “Dream-RSI: Recursive Self-Improvement through Evolving Worlds,” arXiv preprint 2609.14858, Sep 2026. https://arxiv.org/abs/2609.14858
  31. [31]SIMA Team, Google DeepMind, “SIMA 2: A Generalist Embodied Agent for Virtual Worlds,” arXiv preprint 2512.04797, Nov 2025; announcement blog Nov 13, 2025. https://arxiv.org/abs/2512.04797 · https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds
  32. [32]Google DeepMind, “Genie 3: A New Frontier for World Models,” Google DeepMind blog, Aug 2025. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/
  33. [33]Anthropic, “Natural Emergent Misalignment from Reward Hacking in Production RL,” arXiv preprint 2511.18397, Nov 2025. https://arxiv.org/abs/2511.18397
  34. [34]Mechanize, company website (product description), accessed Sep 2026. https://www.mechanize.work/
  35. [35]K. Roose, “This A.I. Company Wants to Take Your Job,” The New York Times, June 11, 2025. https://www.nytimes.com/2025/06/11/technology/ai-mechanize-jobs.html
  36. [36]Business Insider, “Google is in talks for a $1.5 billion-plus deal with AI coding agent startup Mechanize,” Aug 5, 2026. https://www.businessinsider.com/google-is-in-talks-for-a-dollar15-billion-plus-deal-with-ai-coding-agent-startup-mechanize ; Business Insider, “Google Completed Its Talent Deal for AI Agents Startup Mechanize,” September 2026. https://www.businessinsider.com/google-completes-deal-for-ai-agents-startup-mechanize-2026-9
  37. [37]Bespoke Labs, “Bespoke Labs Raises $40M to Build Environments that Enable Reliable Agents,” Jul 6, 2026. https://bespokelabs.ai/blog/bespoke-labs-raises-40m-to-build-environments-that-enable-reliable-agents ; M. Deutscher, “AI post-training startup Bespoke Labs raises $40M in funding,” SiliconANGLE, Jul 6, 2026. https://siliconangle.com/2026/07/06/ai-post-training-startup-bespoke-labs-raises-40m-funding
  38. [38]Prime Intellect, “Environments Hub,” Prime Intellect blog, Aug 27, 2025; environment count from company homepage, Sep 2026. https://www.primeintellect.ai/blog/environments
  39. [39]Mercor, “Mercor to acquire Deeptune,” Mercor blog, Jul 9, 2026. https://www.mercor.com/blog/mercor-to-acquire-deeptune/
  40. [40]PR Newswire, “Realset AI and Flatkey raise $10M Series A to build real-world training data for frontier models and embodied agents,” Sep 23, 2026. https://www.prnewswire.com/news-releases/realset-ai-and-flatkey-raise-10m-series-a-to-build-real-world-training-data-for-frontier-models-and-embodied-agents-302887825.html
  41. [41]Wing VC, “Who Will Win the RL Environment Market (and Why),” Jan 2026 (secondary analysis). https://www.wing.vc/content/who-will-win-the-rl-environment-market--and-why
  42. [42]Office of Sen. Bernie Sanders, “Sanders, Casar introduce legislation to create new federal agency to ban artificial superintelligence, pause advanced AI development,” press release, Sep 23, 2026. https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-create-new-federal-agency-to-ban-artificial-superintelligence-pause-advanced-ai-development/