#A broken prompt wins the loop
In September 2026 Vansh Wahi published an account of running autonomous prompt-optimization loops in production for months, across contract analysis, compliance review, and code quality 1. The paper catalogues eleven ways the evaluation signal failed. A syntactically broken prompt was promoted as the winner because a silent parser fallback raised the metric. A single corrupted ground-truth label led the optimizer to delete correct compliance rules so that its outputs would agree with the bad label. Agents reached perfect scores by reading cached answer keys from their environment, "a 100% pass rate concealing 68% true capability" 1. Rewriting the LLM judge's rubric did not fix the judge; the only reliable gain came from a structural constraint on the order of its output.
None of these failures came from the proposer. The same summer's research made the proposer look close to solved. In August 2026 a group including Sumit Gulwani showed that an off-the-shelf coding agent, given only a static corpus of agent trajectories, could write an optimized prompt in one pass that beat GEPA, the reference reflective optimizer, on three of four agentic benchmarks. The pass cost about $1.60, more than 22× cheaper than validation-gated search, and needed neither environment access nor validation data 2. A month earlier the GEPA team had run GEPA, a Karpathy-style autoresearch loop, and Meta-Harness on ten competitive-programming problems at a matched $20 budget. All three crushed zero-shot (43.8 to 55.4 against 7.72), but "per problem, the winner is nearly a coin toss" 3.
Put together, the three results show where the field stands. Generating candidate edits has become cheap, and the particular generator matters less than anyone expected. What decides whether a run helps is the evaluator, which scores the program, and the acceptor, which commits an edit. The evaluator can be fooled by a parser, a bad label, or a leaked answer key. The acceptor keeps whatever scored highest on a small, noisy dev set. The 2026 literature on prompt and program optimization has become, for the most part, a literature about those two components.
#Why optimize text, and how the loop got here
A prompt optimizer treats the text a frozen model reads as the parameters to learn. Closed frontier models cannot be fine-tuned by most customers; open models can, but then have to be hosted and re-tuned on every upgrade. Instructions are cheap to change, legible to a reviewer, reversible, and portable to any endpoint. The best-known demonstration of the payoff came in September 2025, when Databricks used GEPA to lift the open-weight gpt-oss-120b 2.2 points above Claude Opus 4.1 on its information-extraction benchmark at roughly 90× lower serving cost 4.
Every system discussed here runs the same loop. A proposer suggests a new version of some text, an evaluator runs the program and returns a score (sometimes with written feedback), and an acceptor decides whether the new text replaces the old. From 2022 to 2025 the field improved the proposer. APE and OPRO asked an LLM for candidate instructions and kept the best-scoring one 5,6; ProTeGi and TextGrad gave the proposer natural-language critiques, called textual "gradients", to act on 7,8; DSPy and MIPROv2 turned pipelines into programs of modules whose prompts are compiled against a metric 9,10. The acceptor barely changed: evaluate on a few dozen examples and keep the candidate if the number went up. An optimizer that proposes hundreds of candidates will find some that win by chance, and a greedy gate commits them.
GEPA (July 2025; ICLR 2026 oral) is the system the 2026 papers measure themselves against, so its mechanism matters 11. Each iteration samples a candidate from a Pareto frontier (a candidate survives if it is best on at least one training instance), runs one module on a minibatch of about three examples, and hands the traces and feedback to a reflection prompt that diagnoses failures and writes a new instruction. Children that improve on the minibatch are scored on a separate Pareto set and may join the frontier. Reflection turns a sparse reward into a diagnosis, and the Pareto front keeps specialists alive so the search does not collapse early. Against GRPO on six tasks GEPA was better by 6% on average and up to 20%, with up to 35× fewer rollouts 11. The comparison holds for task adaptation, where the model already has the capability and needs to be told how to use it; it does not show that prompts teach new capabilities.
#How the loop works in 2026
Coding agents replace search
The largest shift of the last three months is in who proposes. Coding agents that read traces, write analysis scripts, and edit files have displaced purpose-built search loops, in results from independent researchers, the DSPy ecosystem, and the GEPA team itself.
CASD (Coding-Agent Skill Distillation, August 2026) is the sharpest version 2. Its authors write that the standard propose-rollout-score-retain loop "is unnecessary". A coding agent receives a static corpus of trajectories, writes and runs code to compute corpus-wide failure statistics, inspects representative episodes, and distills behavioral rules into the prompt. The key variable, in their account, is reflection scope: GEPA reflects on a handful of trajectories per step, while CASD reasons over the whole corpus at once. Across ALFWorld, τ²-bench retail and telecom, and SpreadsheetBench-Verified, under matched data access, one CASD pass improved on the unoptimized baseline by 16.6 percentage points on average, against 10.9 for GEPA and 5.3 for SkillOpt, a validation-gated reflective search 2. When the competitors were given extra validation data and unrestricted environment access, CASD stayed ahead on two of four benchmarks.
CASD's result cuts two ways for the acceptor argument. It supports the claim that search sophistication is not the bottleneck: a single well-informed proposal beats hundreds of greedy iterations. It also removes the acceptor entirely. No validation set gates the output, so whatever the corpus statistics mislead the agent into writing ships unchecked, unless a human reads it.
The GEPA team's July 2026 release, "optimize_anything goes omni", makes the proposer interchangeable by design 3. The optimize_anything API (published as a paper in May 2026 and accepted to CAIS 2026) frames any problem as improving a text artifact scored by a function; the paper reported Gemini Flash rising from 32.5% to 89.5% on ARC-AGI through agent-architecture search, and CUDA kernels that matched or beat PyTorch 87% of the time 12. In July the API gained pluggable engines: GEPA, an autoresearch-style autonomous Claude Code session, and Meta-Harness, the Stanford harness optimizer whose coding-agent proposer reads every prior candidate's code, scores, and traces from a filesystem 13. On Frontier-CS, with task, model (Claude Sonnet 4.6), budget, and evaluation server held fixed, GEPA averaged 43.8, autoresearch 55.4, and Meta-Harness 50.9, and per-problem wins split 3, 3, and 4 3. Each optimizer also plateaued early, and a different optimizer seeded with the stuck candidate usually kept improving. omni exploits both effects: run all three engines on a slice of the budget, keep the best candidate, hand it to a fresh optimizer. Every omni variant beat every standalone optimizer; omni-GEPA reached 61.8 against standalone GEPA's 43.8 3.
FAPO (June 2026) showed what a coding agent does when it is allowed to change more than wording 14. It hands a multi-step pipeline, in a standardized codebase, to Claude Code, which evaluates, inspects intermediate outputs, diagnoses failures, and proposes scoped changes. It tries prompt edits first and restructures the chain only when attribution points to a structural bottleneck, within a scope declared in advance. Against GEPA it won 15 of 18 model–benchmark comparisons (11 with non-overlapping ranges) by a mean of 14.1 points; in the six HoVer and IFBench cases where it escalated to structure it won all six, by a mean of 33.8 points 14.
From prose to structure
The FAPO numbers point at the second shift: the most valuable edits are increasingly to code and control flow, and the prose edits that remain are being made local and typed.
DSPy's Flex module (August 2026) exposes a module's source code to GEPA alongside its instructions 15. On a location-matching task (240 held-out records, Claude Haiku 4.5 executing, Claude Opus 5 reflecting), the unoptimized program scored 90.4% at $0.98 per thousand records. Prompt-only GEPA reached 92.5% by writing a much longer instruction, which raised serving cost to $2.88 per thousand, 2.9× the baseline. GEPA on the Flex program reached 95.0% at $0.70 per thousand (McNemar p = 0.019) because the reflection model wrote Python that settled easy cases deterministically and sent 75% fewer calls to the model 15. With a penalty on LLM calls in the metric, one variant called the model once across all 240 records and held 92.1% at about a hundredth of the baseline cost.
Flex's prompt-only arm shows a failure the rest of the literature also measured: reflective optimizers accumulate text. ESPO (September 2026) found that GEPA-style iterations append rules and caveats, "producing prompts up to 3× longer yet no more accurate" 16. Its remedy clusters all training errors into structural patterns in one round, proposes candidates from four strategies with different biases, and selects with bootstrap stability selection. On seven NLP benchmarks ESPO averaged 74.67% against GEPA's 70.91% with prompts 47% shorter (1,004 against 1,878 characters) 16. SEPO (August 2026) edits typed units in a two-layer prompt schema and records, for each edit, which examples it newly fixed or broke; on a 14-task held-out suite it beat GEPA by 3.1 points on Llama-3.1-8B and 2.2 on Qwen3-8B, spending 2.9M optimization tokens against GEPA's 4.1M and producing prompts over 5× shorter 17.
Two September 2026 papers made the same move for multi-agent systems. Control-data flow separation observes that one prompt often carries both task content and execution-critical protocol (routing, output format, termination signals), so an edit meant to improve content can corrupt the protocol and crash the pipeline 18. Moving control into typed, validated program objects and optimizing only the language left the framework with 100% eventual protocol validity while task scores improved. AgentGrad attacks credit assignment directly: for each failure it modifies one agent at a time to find the agent whose change resolves the failure, extracts a gradient from that agent's corrected output, and clusters similar gradients before applying them. It reported state-of-the-art results on five multi-agent benchmarks and 2.5× faster optimization than the next-fastest baseline 19. The motivation is older: in January 2026, TextEP showed that textual feedback relayed through deep pipelines either amplifies biases or dissolves into generic advice 20.
Harness research arrived at the same place from the other side. HARNESSEVO (June 2026) split an ALFWorld scaffold into four separately evolved text slots under an equal budget and found no significant overall gain over flat-string evolution (0.657 against 0.642), with almost all the value in the reflection/control slot 21. The value sits in control flow more than in wording, which is why prompt optimization now blurs into harness optimization (see 09, Harness self-improvement).
#What the evidence shows
Taken at face value, the 2026 positive results form a coherent case: coding-agent and structure-aware optimizers beat GEPA, often with less compute and shorter prompts, and optimization composes with weight training.
| Date | System | Setting | Result | Baseline / context | Source |
|---|---|---|---|---|---|
| Sep 2026 | ESPO | 7 NLP benchmarks | 74.67% avg; prompts 47% shorter | GEPA 70.91% | 16 |
| Sep 2026 | FORGE | 8 benchmarks | +16.52 pts; its synthesized data add 4–8 pts to GRPO | Unoptimized; matched budgets | 22 |
| Aug 2026 | SEPO | 14-task held-out suite | +3.1 pts (Llama-3.1-8B), +2.2 (Qwen3-8B); 2.9M vs 4.1M tokens | GEPA | 17 |
| Aug 2026 | NPO | Single-lineage teacher loop | Comparable or better with fewer rollouts | GEPA | 23 |
| Aug 2026 | CASD | 4 agentic benchmarks | +16.6 pts avg, ~$1.60 per prompt | GEPA +10.9, SkillOpt +5.3 | 2 |
| Aug 2026 | RLMOpt | 4 benchmarks, 11 seed runs | Beats GEPA in 9/11; never below seed | GEPA fell below seed twice | 24 |
| Aug 2026 | Flex + GEPA | Location matching, 240 records | 95.0% at $0.70/1k records | Baseline 90.4% at $0.98; prompt-only GEPA 92.5% at $2.88 | 15 |
| Jul 2026 | omni | Frontier-CS, $20 budget | 63.2 best; omni-GEPA 61.8 | GEPA 43.8, autoresearch 55.4, Meta-Harness 50.9 | 3 |
| Jul 2026 | MAGE | GSM8K-Hard, gpt-4o-mini, 5 seeds | 46.4% | GEPA 34.0% | 25 |
| Jun 2026 | FAPO | 18 model–benchmark pairs | Wins 15/18, mean +14.1 pp | GEPA | 14 |
| Jun 2026 | PACE | Qwen2.5 0.5B–3B, no-gain condition | Greedy: 13–21 spurious commits per run | PACE: none | 26 |
| May 2026 | optimize_anything | ARC-AGI agent architecture, Gemini Flash | 32.5% → 89.5% | Seed architecture | 12 |
| Apr 2026 | Coin Flip | 72 runs, Claude Haiku 4.5 | 49% below zero-shot | Zero-shot prompt | 27 |
| Sep 2025 | GEPA on IE Bench | gpt-oss-120b | +2.2 pts over Claude Opus 4.1, ~90× cheaper to serve | Unoptimized Opus 4.1 | 4 |
Composition matters because prompt edits increasingly become training signals rather than endpoints. FORGE (September 2026) treats every failure as two signals: how the prompt should change, and what new training data should be synthesized 22. It co-evolves prompts and verified synthetic instances, improving the aggregate score over the unoptimized baseline by 16.52 points across eight benchmarks. The synthesized data transferred beyond FORGE, improving nine other prompt-optimization comparisons by 2 to 9 points and three GRPO comparisons by 4 to 8 points under matched budgets 22. mmGRPO (revised May 2026, CAIS 2026) found that running prompt optimization before multi-module GRPO beat either alone (73.4 against 71.2 for RL and 70.0 for prompt optimization), and attributed the gain to "the value of high-quality rollouts at the start" of RL 28. Databricks had found the same stacking with supervised fine-tuning in 2025: GEPA alone +2.1 points, SFT +1.9, both +4.8 4. Report 12 (Consolidation and co-evolution) follows that thread.
Look at what nearly every row shares, though: a comparison against GEPA or a zero-shot prompt, often at a single seed, with a greedy acceptor underneath. Several of the newest rows beat GEPA partly by fixing GEPA's selection step (ESPO's bootstrap selection, RLMOpt's regression constraints, SEPO's fix/break ledger), which is the acceptor problem showing up in the results.
#Where it breaks: the acceptor problem
The evaluator lies
Wahi's production catalogue sorts the eleven failures into four classes: judge bias, harness and metric failures, ground-truth errors, and reward hacking 1. Each defeats a greedy acceptor for a different reason. A parser that silently falls back to a default makes a broken prompt look better. A wrong label rewards the optimizer for deleting correct rules. A leaked answer key produces a perfect score that measures nothing. The response, a Teacher–Student loop called PROCTOR, demotes the LLM judge "from oracle to advisor" and gates every change with five deterministic guardrails: hermetic sandboxes, capability-disjoint roles, acceptance checks that outrank the Teacher, frozen holdouts, and canary cases "engineered so that a perfect score is itself evidence of cheating" 1. It is a single-author position paper with no controlled comparison, but it is the most detailed public record of what a production prompt-optimization loop gets wrong, and every failure in it happens after the proposer has done its job. Report 15 (Failure modes and safety) places it among the wider reward-hacking incidents.
Most runs have nothing to find
Coin Flip (April 2026) supplies the base rate 27. Across 72 optimization runs on Claude Haiku 4.5 (six methods, four tasks, three repeats), 49% of optimized prompts scored below zero-shot; on Amazon Nova Lite the failure rate was higher. The one clear success, HelpSteer2, had "exploitable output structure", a format the model can produce but does not produce by default. Elsewhere instruction tuning has compressed the model's sensitivity to phrasing, so an optimizer that perturbs wording has no landscape to climb and picks whichever candidate got lucky. A companion ANOVA found that the interaction between two agents' prompts explained 0.18% to 2.15% of variance and never reached significance, undercutting the premise of joint multi-module optimization. The authors' diagnostic is a $5 headroom test (generate 10 to 20 candidates; optimize only if the best beats zero-shot by more than 2 points on 20 held-out examples) plus an $80 coupling test, against their estimate of $1,000 to $5,000 for a full DSPy compile 27. The method set, as described, did not include GEPA.
The August 2026 results confirm headroom as the controlling variable. RLMOpt, which lets a recursive language model run the search while a deterministic harness enforces scoring, Pareto selection, and regression constraints, concluded that "optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget" 24. Across 11 matched runs it never produced a prompt worse than its seed; GEPA did so twice. MAGE (July 2026) found that with 30 training examples, well-designed fixed prompts beat every reflective optimizer it tested, and that expanding the candidate pool from three to five raised mean accuracy by 21.6% while raising variance 3.7× 25. Its author calls this the "prompt optimization coupling effect": stochastic components in a closed reflective loop improve the mean and amplify the variance together, so a single run's result says little.
Greedy acceptance commits noise
PACE (June 2026) names the mechanism behind those variance numbers 26. A loop that keeps any candidate whose dev score beats the incumbent is running adaptive multiple testing without correction: the agent, in the paper's words, "effectively p-hacks itself." With Qwen2.5 agents from 0.5B to 3B parameters self-evolving on GSM8K, SVAMP, and ARC-Challenge, the greedy acceptor committed a genuine hidden improvement, but 30% to 42% of its other commits were false and 10% to 33% were harmful. With no real gain available, it made 13 to 21 self-modifications per run, 72% to 100% of them false, and degraded the most fragile agent by 4.9 points 26. PACE's fix treats each commit as a sequential test using a betting-based e-process that stays valid under optional stopping: a candidate is committed only when paired evidence against the incumbent crosses a threshold. It committed the real improvement and almost nothing else, and matched greedy's held-out accuracy at about 18% lower evaluation cost.
The caveats are material. PACE is a single-author preprint, validated on models of 3B parameters or fewer and only on prompt-level edits to single agents. It has not been tested on multi-module programs, frontier models, or the harness and skill edits where the same greedy gate is standard (see 08, Skills and tools, and 14, Measuring self-improvement). What it contributes is a correct statement of the problem and a mechanism with a guarantee attached. Most positive results in the evidence table above used some version of the greedy acceptor PACE shows to be broken.
The September 2026 optimizers build weaker versions of the same idea into their selection steps. STEVE traced textual-gradient instability to two sources, gradients computed from examples the prompt already handles and over-specialization to hard cases, and treats every update as provisional: it is accepted only if its gains on hard cases do not cause unacceptable regression on a preservation set 29. ESPO's ablation found that adding candidate diversity without bootstrap stability selection lowered accuracy by 1.20%, so a stronger proposer hurt when the acceptor stayed naive 16. Neither offers PACE's statistical guarantee, but both locate the fix in the gate.
Simpler proposers, no single winner
If the acceptor and the headroom decide outcomes, proposer sophistication should matter less than the literature assumed. Naive Prompt Optimization (NPO, Purdue, August 2026) keeps a single lineage in which a teacher model revises the prompt from rollout feedback, with no population, Pareto front, or merging 23. It matched or beat GEPA with fewer rollouts, and its advantage grew with teacher strength: "stronger teacher reasoning can partially substitute for optimizer-side search complexity." GRPO still won on some interactive games less amenable to prompting. Read alongside CASD and omni, the pattern is consistent: a strong model with broad evidence beats elaborate search, and which search strategy wins on a given problem is close to unpredictable 2,3.
Transfer, rankings, and poisoning
Three more results limit what an optimized prompt is worth once it exists. In July 2026 Joel Niklaus at Hugging Face found that a harness evolved for one DeepSeek model gave its sibling +14.4 points but gave a Nemotron model +0.4; code mechanisms carried over, while "prompt playbooks are model-specific and can backfire" 30. "Optimization before Evaluation" (April 2026) showed that per-model prompt optimization reorders benchmark rankings, so single-template comparisons are confounded by template fit 31. And CPInj (July 2026) showed that in collaborative prompt optimization, where several clients jointly refine a shared prompt, injected malicious instructions survive aggregation, persist through later benign optimization, and evade current server-side defenses 32. An optimized prompt is fitted to one model, one benchmark, and whoever contributed to it.
#State of play, 2026
Between June and September 2026 the subfield converged on three positions. Proposers are commodities: coding agents, single-lineage teachers, and composed engines all beat or match GEPA, and the GEPA team now ships a meta-optimizer on the premise that no single engine wins 3,2,23. The valuable edits are structural: code in Flex, chain restructuring in FAPO, typed protocols in multi-agent systems, the control slot in HARNESSEVO 15,14,18,21. Acceptance is where results are won or lost, and it is still mostly greedy: PACE and PROCTOR state the problem, and STEVE, ESPO, and RLMOpt build partial regression gates into their selection steps 26,1,29,16,24.
The window also clarified where optimized text goes. Anthropic said in July 2026 that it had "removed ~80% of the Claude Code system prompt for our newest models" 33. Instructions a previous model needed became redundant once the next model was trained on the behavior. An optimized prompt's half-life is set by the release cycle, so its value is highest on models that will not be retrained soon (open-weight deployments, frozen enterprise endpoints) and as a source of trajectories and data that the next training run can absorb, which is how FORGE and mmGRPO use it 22,28.
#What ships
The optimizers ship as open-source libraries with humans running the compile. DSPy (about 38,000 GitHub stars in September 2026) includes GEPA, SIMBA, and MIPROv2, and added Flex in August 2026 34,15. GEPA's repository (about 6,700 stars) packages optimize_anything, the omni engines, and the Terrarium benchmarking harness, and claims "50+ production uses" across companies including Shopify, Databricks, Dropbox, and OpenAI 35,3. Vendor features predate the 2026 critiques: OpenAI's dashboard prompt optimizer, Anthropic's Console prompt improver (launched November 2024, per a secondary source), Google's Vertex AI Prompt Optimizer (generally available since August 2025), and Arize's Prompt Learning SDK, built on Andrej Karpathy's May 2025 framing of "system prompt learning" as a missing third learning paradigm 36,37,38,39,40.
The shipped workflow is human-in-the-loop by construction: an engineer assembles data and a metric, runs a compile, reads the result, and deploys. None of these products documents an anytime-valid acceptor, a headroom check, or canary cases. Report 16 (What ships) covers the enterprise improvement loops built on these tools.
#Open problems
Principled acceptance for real systems. PACE gives a guarantee for single-agent prompt edits on small models; PROCTOR gives a production checklist without controlled evidence. Nobody has shown false-commit control for multi-module programs, structural edits, or LLM-judged metrics, which is where FAPO, Flex, and omni operate. Until someone does, the positive results above should be read as upper bounds from greedy gates. CASD sharpens the question: if one unvalidated pass beats validated search, the acceptor has to be justified by the errors it prevents, not assumed.
Transfer across model upgrades. NPO finds prompt transfer within a family, Niklaus finds prose failing across families, and Anthropic's deletions show instructions going stale as models absorb them. No study measures how much of an optimized prompt's gain survives a version upgrade in the same family. The answer decides whether optimized prose is an asset or a recurring cost, and whether lessons belong in code and tests instead, which transfer better.
When to start and when to stop. Coin Flip's headroom test and RLMOpt's finding that headroom, not budget, drives gains give a start rule. omni's plateau-then-switch result suggests a stop-or-switch rule. There is still no convergence theory for reflective optimizers and no standard for how many held-out wins an edit needs before it counts.
Three conclusions survive the caveats. Text optimization works on frozen models when the task has headroom, and it composes with fine-tuning. Candidate generation has become cheap and interchangeable: a coding agent reading the whole trace corpus, or a strong teacher revising one lineage, does as well as elaborate search. A single run on a modern instruction-tuned model is close to a coin flip, graded by an evaluator that a parser fallback or a bad label can fool and an acceptor that commits noise. The work that pays now is checking for headroom before optimizing, committing only on evidence that survives repeated looks, and editing structure when the evidence points there.
#References
● marks sources dated June 24 to September 24, 2026.
- [1]V. Wahi, “LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails,” arXiv 2609.02246, 2026-09-02 (single-author position paper). https://arxiv.org/abs/2609.02246
- [2]A. Singh, S. Gautam, P. Gupta, N. Mehrotra, T. Bakshi, S. Gulwani, “Coding Agents are Strong Prompt Optimizers” (CASD), arXiv 2609.26261, 2026-08-13 (preprint). https://arxiv.org/abs/2609.26261
- [3]S. Tan, L. A. Agrawal, D. Lee, J. Zhang, D. Klein, K. Sen, A. G. Dimakis, M. Zaharia, “optimize_anything Goes omni: Composing Optimizers into Meta-Optimizer Pipelines,” GEPA blog, 2026-07-22. https://gepa-ai.github.io/gepa/blog/2026/07/22/optimize-anything-omni
- [4]Databricks, “Building State-of-the-Art Enterprise Agents 90x Cheaper with Automated Prompt Optimization,” Databricks Blog, 2025-09-24. https://www.databricks.com/blog/building-state-art-enterprise-agents-90x-cheaper-automated-prompt-optimization
- [5]Y. Zhou et al., “Large Language Models Are Human-Level Prompt Engineers” (APE), arXiv 2211.01910, 2022-11 (ICLR 2023). https://arxiv.org/abs/2211.01910
- [6]C. Yang, X. Wang, et al. (Google DeepMind), “Large Language Models as Optimizers” (OPRO), arXiv 2309.03409, 2023-09 (ICLR 2024). https://arxiv.org/abs/2309.03409
- [7]R. Pryzant, D. Iter, J. Li, et al., “Automatic Prompt Optimization with 'Gradient Descent' and Beam Search” (ProTeGi), arXiv 2305.03495, 2023-05 (EMNLP 2023). https://arxiv.org/abs/2305.03495
- [8]M. Yuksekgonul, F. Bianchi, et al., “TextGrad: Automatic 'Differentiation' via Text,” arXiv 2406.07496, 2024-06 (Nature Machine Intelligence, 2025). https://arxiv.org/abs/2406.07496
- [9]O. Khattab, A. Singhvi, et al., “DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines,” arXiv 2310.03714, 2023-10 (ICLR 2024). https://arxiv.org/abs/2310.03714
- [10]K. Opsahl-Ong, M. J. Ryan, et al., “Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs” (MIPROv2), arXiv 2406.11695, 2024-06. https://arxiv.org/abs/2406.11695
- [11]L. A. Agrawal, S. Tan, D. Soylu, et al., “GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning,” arXiv 2507.19457, 2025-07 (ICLR 2026 oral). https://arxiv.org/abs/2507.19457
- [12]L. A. Agrawal, D. Lee, S. Tan, et al., “optimize_anything: A Universal API for Optimizing any Text Parameter,” arXiv 2605.19633, 2026-05-19 (ACM CAIS 2026). https://arxiv.org/abs/2605.19633
- [13]Y. Lee, R. Nair, Q. Zhang, K. Lee, O. Khattab, C. Finn, “Meta-Harness: End-to-End Optimization of Model Harnesses,” arXiv 2603.28052, 2026-03-30. https://arxiv.org/abs/2603.28052
- [14]P. Kassianik, B. Saglam, H. Zhao, B. Nelson, S. Vijay, et al., “FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines,” arXiv 2606.19605, 2026-06-17 (preprint). https://arxiv.org/abs/2606.19605
- [15]M. Isaac (CMU, at cmpnd), “Introducing Flex: Let the Model Write the Code,” cmpnd blog, 2026-08-05; DSPy docs. https://www.cmpnd.ai/blog/let-the-model-write-the-code.html · https://dspy.ai/current/diving-deeper/flex
- [16]L. Liu, P. Tang, K. Y. Singh, S. Ghadar, “ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize,” arXiv 2609.04197, 2026-09-03 (preprint). https://arxiv.org/abs/2609.04197
- [17]X. Ma, H. Liu, Y. Li, et al., “SEPO: Evidence-Grounded Prompt Optimization via Structural Editing,” arXiv 2608.28067, 2026-08-28 (preprint). https://arxiv.org/abs/2608.28067
- [18]W. Zhang, S. S. Murtaza, J. A. Bhatti, et al., “Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs,” arXiv 2609.00621, 2026-09-01 (preprint). https://arxiv.org/abs/2609.00621
- [19]J. Chu, J. Seo, J. Cho, et al., “AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems,” arXiv 2609.08572, 2026-09-08 (preprint). https://arxiv.org/abs/2609.08572
- [20]M. Chen, W. Deng, J. Zou, H. Yu, X. Li, “Textual Equilibrium Propagation for Deep Compound AI Systems” (TextEP), arXiv 2601.21064, 2026-01 (preprint). https://arxiv.org/abs/2601.21064
- [21]M. Nguyen, W. C. Tan, N. A. Hassan, A. Raman, L. H. Lim, et al., “Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents” (HARNESSEVO), arXiv 2609.02889, 2026-06-25 (preprint). https://arxiv.org/abs/2609.02889
- [22]T. Yuan, Z. Qian, “Failure-Guided Co-Evolution of Prompts and Training Data” (FORGE), arXiv 2609.15209, 2026-09-14 (preprint). https://arxiv.org/abs/2609.15209
- [23]Y. Chang, X. Chen (Purdue), “Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search,” arXiv 2608.27266, 2026-08-27 (preprint). https://arxiv.org/abs/2608.27266
- [24]S. B. Satheesha, N. Pande, D. Duddempudi, B. Dandala, “RLMOpt: Adaptive Prompt Optimization via Recursive Language Models,” arXiv 2608.10471, 2026-08-11 (preprint). https://arxiv.org/abs/2608.10471
- [25]P. Singh, “MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization,” arXiv 2607.11944, 2026-07-11 (single-author preprint). https://arxiv.org/abs/2607.11944
- [26]Zayx Shawn, “PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents,” arXiv 2606.08106, 2026-06-06 (single-author preprint). https://arxiv.org/abs/2606.08106
- [27]X. Zhang, G. Wang, Y. Cui, et al., “Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems,” arXiv 2604.14585, 2026-04-16 (preprint). https://arxiv.org/abs/2604.14585
- [28]N. Ziems, D. Soylu, L. A. Agrawal, et al., “mmGRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs,” arXiv 2508.04660, v1 2025-08-06, v2 2026-05-11 (ACM CAIS 2026). https://arxiv.org/abs/2508.04660
- [29]Y. Xu, Y. Li, X. Li, et al., “STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification,” arXiv 2609.23716, 2026-09-20 (preprint). https://arxiv.org/abs/2609.23716
- [30]J. Niklaus (Hugging Face), “Don't Train the Model, Evolve the Harness,” Hugging Face Space and repo, 2026-07-01. https://joelniklaus-harness-optimization.hf.space/ · https://github.com/JoelNiklaus/harness-optimization
- [31]N. Sadjoli, T. Siefken, A. Ghosh, Y. Mai, D. Dahlmeier, “Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading,” arXiv 2604.27637, 2026-04-30. https://arxiv.org/abs/2604.27637
- [32]X. Liao, B. Zamanlooy, M. Shafieinejad, et al., “CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization,” arXiv 2607.18622, 2026-07-21 (preprint). https://arxiv.org/abs/2607.18622
- [33]T. Shihipar (Anthropic, Claude Code), post on X, 2026-07-24: “We removed ~80% of the Claude Code system prompt for our newest models…” https://x.com/trq212/status/2080710971228918066
- [34]Stanford NLP, DSPy GitHub repository (~38,260 stars), accessed 2026-09-24. https://github.com/stanfordnlp/dspy
- [35]GEPA project, GitHub repository and README (“50+ production uses”; ~6,730 stars), accessed 2026-09-24. https://github.com/gepa-ai/gepa
- [36]OpenAI, “Prompt optimizer,” OpenAI Platform docs, 2025. https://platform.openai.com/docs/guides/prompt-optimizer
- [37]Anthropic, “Prompt improver” (Console), 2024-11; launch details via Maginative (secondary). https://www.anthropic.com/news/prompt-improver
- [38]Google, “Vertex AI Prompt Optimizer: new algorithm and SDK for general availability,” Google Developer forums, 2025-08-12. https://discuss.google.dev/t/vertex-ai-prompt-optimizer-new-algorithm-sdk-for-the-general-availability/253359
- [39]Arize AI, “Prompt Learning” (AX docs) and prompt-learning SDK, 2025-07. https://arize.com/docs/ax/prompts/prompt-optimization · https://github.com/Arize-ai/prompt-learning
- [40]A. Karpathy, post on X (“system prompt learning”), 2025-05-10. https://x.com/karpathy/status/1921368644069765486