Research series · 2026 · Report 06 of 16
Report 06 of 16 · Part II · Non-parametric · Prompts and programs

Prompt and program optimization

from better prompts to provable edits

By mid-2026 proposing better prompts had become cheap and interchangeable: a coding agent's single pass beats GEPA, and the GEPA team's own study finds the winning optimizer changes from problem to problem. Deciding which edits are real is still expensive and fragile. Nearly half of optimization runs fail to beat zero-shot, the standard "keep it if the score went up" acceptor commits false edits, and production loops promote broken prompts when a metric misfires. The work that matters now is about acceptance and about structure: gating edits on evidence, and editing code and control flow instead of prose.

ContentsReport 06 · Non-parametric
$ tree ./06-prompt-program-optimization
./06-prompt-program-optimization
├── 01-a-broken-prompt-wins-the-loop# 341 words
├── 02-why-optimize-text-and-how-the-loop-got…# 389 words · 1 figure
├── 03-how-the-loop-works-in-2026# 1064 words · 1 figure
├── 04-what-the-evidence-shows# 471 words
├── 05-where-it-breaks-the-acceptor-problem# 1094 words
├── 06-state-of-play-2026# 197 words
├── 07-what-ships# 163 words
└── 08-open-problems# 320 words
8 sections · 40 references · 2 figures
~$1.60
one coding-agent pass, 22× cheaper than gated search
30–42%
false commits under greedy acceptance
49%
of 72 optimization runs scored below zero-shot
40
references · 19 from Jun 24 – Sep 24, 2026
17 min
reading time
Storage
Non-parametric: instructions, demonstrations, system prompts, multi-module program text; by 2026 also module code, kernels, agent architectures
Engines
LLM reflection over traces; Pareto and evolutionary search; coding agents as optimizers; composed meta-optimizers
Evaluator
Task metric on a dev set, often with LLM-written feedback; execution for code artifacts; LLM judges in production
Loop timescale
Offline batch compiles, minutes to hours per run
Loop closure
Human-in-the-loop: an engineer runs the compile and decides what ships
Evidence maturity
GEPA and optimize_anything are peer-reviewed (ICLR 2026, CAIS 2026); nearly all June–September 2026 results, including every acceptor critique, are preprints or vendor posts

Where the lesson is written, what writes it, who checks it, how often it runs, whether the loop closes, and how strong the evidence is.

Section 01 / 08

#A broken prompt wins the loop

In September 2026 Vansh Wahi published an account of running autonomous prompt-optimization loops in production for months, across contract analysis, compliance review, and code quality 1. The paper catalogues eleven ways the evaluation signal failed. A syntactically broken prompt was promoted as the winner because a silent parser fallback raised the metric. A single corrupted ground-truth label led the optimizer to delete correct compliance rules so that its outputs would agree with the bad label. Agents reached perfect scores by reading cached answer keys from their environment, "a 100% pass rate concealing 68% true capability" 1. Rewriting the LLM judge's rubric did not fix the judge; the only reliable gain came from a structural constraint on the order of its output.

None of these failures came from the proposer. The same summer's research made the proposer look close to solved. In August 2026 a group including Sumit Gulwani showed that an off-the-shelf coding agent, given only a static corpus of agent trajectories, could write an optimized prompt in one pass that beat GEPA, the reference reflective optimizer, on three of four agentic benchmarks. The pass cost about $1.60, more than 22× cheaper than validation-gated search, and needed neither environment access nor validation data 2. A month earlier the GEPA team had run GEPA, a Karpathy-style autoresearch loop, and Meta-Harness on ten competitive-programming problems at a matched $20 budget. All three crushed zero-shot (43.8 to 55.4 against 7.72), but "per problem, the winner is nearly a coin toss" 3.

Put together, the three results show where the field stands. Generating candidate edits has become cheap, and the particular generator matters less than anyone expected. What decides whether a run helps is the evaluator, which scores the program, and the acceptor, which commits an edit. The evaluator can be fooled by a parser, a bad label, or a leaked answer key. The acceptor keeps whatever scored highest on a small, noisy dev set. The 2026 literature on prompt and program optimization has become, for the most part, a literature about those two components.

Section 02 / 08

#Why optimize text, and how the loop got here

A prompt optimizer treats the text a frozen model reads as the parameters to learn. Closed frontier models cannot be fine-tuned by most customers; open models can, but then have to be hosted and re-tuned on every upgrade. Instructions are cheap to change, legible to a reviewer, reversible, and portable to any endpoint. The best-known demonstration of the payoff came in September 2025, when Databricks used GEPA to lift the open-weight gpt-oss-120b 2.2 points above Claude Opus 4.1 on its information-extraction benchmark at roughly 90× lower serving cost 4.

Every system discussed here runs the same loop. A proposer suggests a new version of some text, an evaluator runs the program and returns a score (sometimes with written feedback), and an acceptor decides whether the new text replaces the old. From 2022 to 2025 the field improved the proposer. APE and OPRO asked an LLM for candidate instructions and kept the best-scoring one 5,6; ProTeGi and TextGrad gave the proposer natural-language critiques, called textual "gradients", to act on 7,8; DSPy and MIPROv2 turned pipelines into programs of modules whose prompts are compiled against a metric 9,10. The acceptor barely changed: evaluate on a few dozen examples and keep the candidate if the number went up. An optimizer that proposes hundreds of candidates will find some that win by chance, and a greedy gate commits them.

Fig 06.1 · Proposer, evaluator, acceptorLoop anatomy · 2026 evidence
COMMIT = NEW INCUMBENT Proposer reflection, search, coding agent Evaluator runs the program, scores it Acceptor keep it if the score went up WHERE IT BREAKS Rarely the bottleneck One coding-agent pass at ~$1.60 per prompt beat GEPA on three of four benchmarks Fooled by the harness silent parser fallback one corrupted label cached answer keys Commits noise With no real gain available: 13 to 21 commits per run, 72% to 100% of them false PACE: commit only when paired evidence crosses a threshold
Proposing is cheap, accepting is not. Every optimizer runs the same three-part loop, and the 2026 evidence places its failures after the proposer. A single coding-agent pass, with no validation data, beat GEPA on three of four benchmarks 2; production evaluators were fooled by a parser fallback, a bad label and leaked answer keys 1; and a greedy acceptor with no real gain to find still committed 13 to 21 edits per run, 72% to 100% of them false, which PACE's sequential test avoids 26.

GEPA (July 2025; ICLR 2026 oral) is the system the 2026 papers measure themselves against, so its mechanism matters 11. Each iteration samples a candidate from a Pareto frontier (a candidate survives if it is best on at least one training instance), runs one module on a minibatch of about three examples, and hands the traces and feedback to a reflection prompt that diagnoses failures and writes a new instruction. Children that improve on the minibatch are scored on a separate Pareto set and may join the frontier. Reflection turns a sparse reward into a diagnosis, and the Pareto front keeps specialists alive so the search does not collapse early. Against GRPO on six tasks GEPA was better by 6% on average and up to 20%, with up to 35× fewer rollouts 11. The comparison holds for task adaptation, where the model already has the capability and needs to be told how to use it; it does not show that prompts teach new capabilities.

Section 03 / 08

#How the loop works in 2026

Coding agents replace search

The largest shift of the last three months is in who proposes. Coding agents that read traces, write analysis scripts, and edit files have displaced purpose-built search loops, in results from independent researchers, the DSPy ecosystem, and the GEPA team itself.

CASD (Coding-Agent Skill Distillation, August 2026) is the sharpest version 2. Its authors write that the standard propose-rollout-score-retain loop "is unnecessary". A coding agent receives a static corpus of trajectories, writes and runs code to compute corpus-wide failure statistics, inspects representative episodes, and distills behavioral rules into the prompt. The key variable, in their account, is reflection scope: GEPA reflects on a handful of trajectories per step, while CASD reasons over the whole corpus at once. Across ALFWorld, τ²-bench retail and telecom, and SpreadsheetBench-Verified, under matched data access, one CASD pass improved on the unoptimized baseline by 16.6 percentage points on average, against 10.9 for GEPA and 5.3 for SkillOpt, a validation-gated reflective search 2. When the competitors were given extra validation data and unrestricted environment access, CASD stayed ahead on two of four benchmarks.

CASD's result cuts two ways for the acceptor argument. It supports the claim that search sophistication is not the bottleneck: a single well-informed proposal beats hundreds of greedy iterations. It also removes the acceptor entirely. No validation set gates the output, so whatever the corpus statistics mislead the agent into writing ships unchecked, unless a human reads it.

The GEPA team's July 2026 release, "optimize_anything goes omni", makes the proposer interchangeable by design 3. The optimize_anything API (published as a paper in May 2026 and accepted to CAIS 2026) frames any problem as improving a text artifact scored by a function; the paper reported Gemini Flash rising from 32.5% to 89.5% on ARC-AGI through agent-architecture search, and CUDA kernels that matched or beat PyTorch 87% of the time 12. In July the API gained pluggable engines: GEPA, an autoresearch-style autonomous Claude Code session, and Meta-Harness, the Stanford harness optimizer whose coding-agent proposer reads every prior candidate's code, scores, and traces from a filesystem 13. On Frontier-CS, with task, model (Claude Sonnet 4.6), budget, and evaluation server held fixed, GEPA averaged 43.8, autoresearch 55.4, and Meta-Harness 50.9, and per-problem wins split 3, 3, and 4 3. Each optimizer also plateaued early, and a different optimizer seeded with the stuck candidate usually kept improving. omni exploits both effects: run all three engines on a slice of the budget, keep the best candidate, hand it to a fresh optimizer. Every omni variant beat every standalone optimizer; omni-GEPA reached 61.8 against standalone GEPA's 43.8 3.

Fig 06.1 · No single optimizer wins Frontier-CSGEPA team · omni · Jul 2026
FRONTIER-CS · AVERAGE SCORE · TEN PROBLEMS · $20 BUDGET PER-PROBLEM WINS 0 Zero-shot 7.72 STANDALONE GEPA 43.8 3 Meta-Harness 50.9 4 Autoresearch 55.4 3 OMNI (COMPOSED) omni-GEPA 61.8 Best omni variant 63.2
Composition beats choosing. With task, model, budget and evaluation server held fixed, the three standalone optimizers all far exceed zero-shot, but per-problem wins split 3, 4 and 3, so no engine is reliably best. Running all three on a slice of the budget and handing the best candidate to a fresh optimizer (omni) beat every standalone run, with omni-GEPA at 61.8 against standalone GEPA's 43.8 3.

FAPO (June 2026) showed what a coding agent does when it is allowed to change more than wording 14. It hands a multi-step pipeline, in a standardized codebase, to Claude Code, which evaluates, inspects intermediate outputs, diagnoses failures, and proposes scoped changes. It tries prompt edits first and restructures the chain only when attribution points to a structural bottleneck, within a scope declared in advance. Against GEPA it won 15 of 18 model–benchmark comparisons (11 with non-overlapping ranges) by a mean of 14.1 points; in the six HoVer and IFBench cases where it escalated to structure it won all six, by a mean of 33.8 points 14.

From prose to structure

The FAPO numbers point at the second shift: the most valuable edits are increasingly to code and control flow, and the prose edits that remain are being made local and typed.

DSPy's Flex module (August 2026) exposes a module's source code to GEPA alongside its instructions 15. On a location-matching task (240 held-out records, Claude Haiku 4.5 executing, Claude Opus 5 reflecting), the unoptimized program scored 90.4% at $0.98 per thousand records. Prompt-only GEPA reached 92.5% by writing a much longer instruction, which raised serving cost to $2.88 per thousand, 2.9× the baseline. GEPA on the Flex program reached 95.0% at $0.70 per thousand (McNemar p = 0.019) because the reflection model wrote Python that settled easy cases deterministically and sent 75% fewer calls to the model 15. With a penalty on LLM calls in the metric, one variant called the model once across all 240 records and held 92.1% at about a hundredth of the baseline cost.

Flex's prompt-only arm shows a failure the rest of the literature also measured: reflective optimizers accumulate text. ESPO (September 2026) found that GEPA-style iterations append rules and caveats, "producing prompts up to 3× longer yet no more accurate" 16. Its remedy clusters all training errors into structural patterns in one round, proposes candidates from four strategies with different biases, and selects with bootstrap stability selection. On seven NLP benchmarks ESPO averaged 74.67% against GEPA's 70.91% with prompts 47% shorter (1,004 against 1,878 characters) 16. SEPO (August 2026) edits typed units in a two-layer prompt schema and records, for each edit, which examples it newly fixed or broke; on a 14-task held-out suite it beat GEPA by 3.1 points on Llama-3.1-8B and 2.2 on Qwen3-8B, spending 2.9M optimization tokens against GEPA's 4.1M and producing prompts over 5× shorter 17.

Two September 2026 papers made the same move for multi-agent systems. Control-data flow separation observes that one prompt often carries both task content and execution-critical protocol (routing, output format, termination signals), so an edit meant to improve content can corrupt the protocol and crash the pipeline 18. Moving control into typed, validated program objects and optimizing only the language left the framework with 100% eventual protocol validity while task scores improved. AgentGrad attacks credit assignment directly: for each failure it modifies one agent at a time to find the agent whose change resolves the failure, extracts a gradient from that agent's corrected output, and clusters similar gradients before applying them. It reported state-of-the-art results on five multi-agent benchmarks and 2.5× faster optimization than the next-fastest baseline 19. The motivation is older: in January 2026, TextEP showed that textual feedback relayed through deep pipelines either amplifies biases or dissolves into generic advice 20.

Harness research arrived at the same place from the other side. HARNESSEVO (June 2026) split an ALFWorld scaffold into four separately evolved text slots under an equal budget and found no significant overall gain over flat-string evolution (0.657 against 0.642), with almost all the value in the reflection/control slot 21. The value sits in control flow more than in wording, which is why prompt optimization now blurs into harness optimization (see 09, Harness self-improvement).

Section 04 / 08

#What the evidence shows

Taken at face value, the 2026 positive results form a coherent case: coding-agent and structure-aware optimizers beat GEPA, often with less compute and shorter prompts, and optimization composes with weight training.

Date System Setting Result Baseline / context Source
Sep 2026 ESPO 7 NLP benchmarks 74.67% avg; prompts 47% shorter GEPA 70.91% 16
Sep 2026 FORGE 8 benchmarks +16.52 pts; its synthesized data add 4–8 pts to GRPO Unoptimized; matched budgets 22
Aug 2026 SEPO 14-task held-out suite +3.1 pts (Llama-3.1-8B), +2.2 (Qwen3-8B); 2.9M vs 4.1M tokens GEPA 17
Aug 2026 NPO Single-lineage teacher loop Comparable or better with fewer rollouts GEPA 23
Aug 2026 CASD 4 agentic benchmarks +16.6 pts avg, ~$1.60 per prompt GEPA +10.9, SkillOpt +5.3 2
Aug 2026 RLMOpt 4 benchmarks, 11 seed runs Beats GEPA in 9/11; never below seed GEPA fell below seed twice 24
Aug 2026 Flex + GEPA Location matching, 240 records 95.0% at $0.70/1k records Baseline 90.4% at $0.98; prompt-only GEPA 92.5% at $2.88 15
Jul 2026 omni Frontier-CS, $20 budget 63.2 best; omni-GEPA 61.8 GEPA 43.8, autoresearch 55.4, Meta-Harness 50.9 3
Jul 2026 MAGE GSM8K-Hard, gpt-4o-mini, 5 seeds 46.4% GEPA 34.0% 25
Jun 2026 FAPO 18 model–benchmark pairs Wins 15/18, mean +14.1 pp GEPA 14
Jun 2026 PACE Qwen2.5 0.5B–3B, no-gain condition Greedy: 13–21 spurious commits per run PACE: none 26
May 2026 optimize_anything ARC-AGI agent architecture, Gemini Flash 32.5% → 89.5% Seed architecture 12
Apr 2026 Coin Flip 72 runs, Claude Haiku 4.5 49% below zero-shot Zero-shot prompt 27
Sep 2025 GEPA on IE Bench gpt-oss-120b +2.2 pts over Claude Opus 4.1, ~90× cheaper to serve Unoptimized Opus 4.1 4

Composition matters because prompt edits increasingly become training signals rather than endpoints. FORGE (September 2026) treats every failure as two signals: how the prompt should change, and what new training data should be synthesized 22. It co-evolves prompts and verified synthetic instances, improving the aggregate score over the unoptimized baseline by 16.52 points across eight benchmarks. The synthesized data transferred beyond FORGE, improving nine other prompt-optimization comparisons by 2 to 9 points and three GRPO comparisons by 4 to 8 points under matched budgets 22. mmGRPO (revised May 2026, CAIS 2026) found that running prompt optimization before multi-module GRPO beat either alone (73.4 against 71.2 for RL and 70.0 for prompt optimization), and attributed the gain to "the value of high-quality rollouts at the start" of RL 28. Databricks had found the same stacking with supervised fine-tuning in 2025: GEPA alone +2.1 points, SFT +1.9, both +4.8 4. Report 12 (Consolidation and co-evolution) follows that thread.

Look at what nearly every row shares, though: a comparison against GEPA or a zero-shot prompt, often at a single seed, with a greedy acceptor underneath. Several of the newest rows beat GEPA partly by fixing GEPA's selection step (ESPO's bootstrap selection, RLMOpt's regression constraints, SEPO's fix/break ledger), which is the acceptor problem showing up in the results.

Section 05 / 08

#Where it breaks: the acceptor problem

The evaluator lies

Wahi's production catalogue sorts the eleven failures into four classes: judge bias, harness and metric failures, ground-truth errors, and reward hacking 1. Each defeats a greedy acceptor for a different reason. A parser that silently falls back to a default makes a broken prompt look better. A wrong label rewards the optimizer for deleting correct rules. A leaked answer key produces a perfect score that measures nothing. The response, a Teacher–Student loop called PROCTOR, demotes the LLM judge "from oracle to advisor" and gates every change with five deterministic guardrails: hermetic sandboxes, capability-disjoint roles, acceptance checks that outrank the Teacher, frozen holdouts, and canary cases "engineered so that a perfect score is itself evidence of cheating" 1. It is a single-author position paper with no controlled comparison, but it is the most detailed public record of what a production prompt-optimization loop gets wrong, and every failure in it happens after the proposer has done its job. Report 15 (Failure modes and safety) places it among the wider reward-hacking incidents.

Most runs have nothing to find

Coin Flip (April 2026) supplies the base rate 27. Across 72 optimization runs on Claude Haiku 4.5 (six methods, four tasks, three repeats), 49% of optimized prompts scored below zero-shot; on Amazon Nova Lite the failure rate was higher. The one clear success, HelpSteer2, had "exploitable output structure", a format the model can produce but does not produce by default. Elsewhere instruction tuning has compressed the model's sensitivity to phrasing, so an optimizer that perturbs wording has no landscape to climb and picks whichever candidate got lucky. A companion ANOVA found that the interaction between two agents' prompts explained 0.18% to 2.15% of variance and never reached significance, undercutting the premise of joint multi-module optimization. The authors' diagnostic is a $5 headroom test (generate 10 to 20 candidates; optimize only if the best beats zero-shot by more than 2 points on 20 held-out examples) plus an $80 coupling test, against their estimate of $1,000 to $5,000 for a full DSPy compile 27. The method set, as described, did not include GEPA.

The August 2026 results confirm headroom as the controlling variable. RLMOpt, which lets a recursive language model run the search while a deterministic harness enforces scoring, Pareto selection, and regression constraints, concluded that "optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget" 24. Across 11 matched runs it never produced a prompt worse than its seed; GEPA did so twice. MAGE (July 2026) found that with 30 training examples, well-designed fixed prompts beat every reflective optimizer it tested, and that expanding the candidate pool from three to five raised mean accuracy by 21.6% while raising variance 3.7× 25. Its author calls this the "prompt optimization coupling effect": stochastic components in a closed reflective loop improve the mean and amplify the variance together, so a single run's result says little.

Greedy acceptance commits noise

PACE (June 2026) names the mechanism behind those variance numbers 26. A loop that keeps any candidate whose dev score beats the incumbent is running adaptive multiple testing without correction: the agent, in the paper's words, "effectively p-hacks itself." With Qwen2.5 agents from 0.5B to 3B parameters self-evolving on GSM8K, SVAMP, and ARC-Challenge, the greedy acceptor committed a genuine hidden improvement, but 30% to 42% of its other commits were false and 10% to 33% were harmful. With no real gain available, it made 13 to 21 self-modifications per run, 72% to 100% of them false, and degraded the most fragile agent by 4.9 points 26. PACE's fix treats each commit as a sequential test using a betting-based e-process that stays valid under optional stopping: a candidate is committed only when paired evidence against the incumbent crosses a threshold. It committed the real improvement and almost nothing else, and matched greedy's held-out accuracy at about 18% lower evaluation cost.

The caveats are material. PACE is a single-author preprint, validated on models of 3B parameters or fewer and only on prompt-level edits to single agents. It has not been tested on multi-module programs, frontier models, or the harness and skill edits where the same greedy gate is standard (see 08, Skills and tools, and 14, Measuring self-improvement). What it contributes is a correct statement of the problem and a mechanism with a guarantee attached. Most positive results in the evidence table above used some version of the greedy acceptor PACE shows to be broken.

The September 2026 optimizers build weaker versions of the same idea into their selection steps. STEVE traced textual-gradient instability to two sources, gradients computed from examples the prompt already handles and over-specialization to hard cases, and treats every update as provisional: it is accepted only if its gains on hard cases do not cause unacceptable regression on a preservation set 29. ESPO's ablation found that adding candidate diversity without bootstrap stability selection lowered accuracy by 1.20%, so a stronger proposer hurt when the acceptor stayed naive 16. Neither offers PACE's statistical guarantee, but both locate the fix in the gate.

Simpler proposers, no single winner

If the acceptor and the headroom decide outcomes, proposer sophistication should matter less than the literature assumed. Naive Prompt Optimization (NPO, Purdue, August 2026) keeps a single lineage in which a teacher model revises the prompt from rollout feedback, with no population, Pareto front, or merging 23. It matched or beat GEPA with fewer rollouts, and its advantage grew with teacher strength: "stronger teacher reasoning can partially substitute for optimizer-side search complexity." GRPO still won on some interactive games less amenable to prompting. Read alongside CASD and omni, the pattern is consistent: a strong model with broad evidence beats elaborate search, and which search strategy wins on a given problem is close to unpredictable 2,3.

Transfer, rankings, and poisoning

Three more results limit what an optimized prompt is worth once it exists. In July 2026 Joel Niklaus at Hugging Face found that a harness evolved for one DeepSeek model gave its sibling +14.4 points but gave a Nemotron model +0.4; code mechanisms carried over, while "prompt playbooks are model-specific and can backfire" 30. "Optimization before Evaluation" (April 2026) showed that per-model prompt optimization reorders benchmark rankings, so single-template comparisons are confounded by template fit 31. And CPInj (July 2026) showed that in collaborative prompt optimization, where several clients jointly refine a shared prompt, injected malicious instructions survive aggregation, persist through later benign optimization, and evade current server-side defenses 32. An optimized prompt is fitted to one model, one benchmark, and whoever contributed to it.

Section 06 / 08

#State of play, 2026

Between June and September 2026 the subfield converged on three positions. Proposers are commodities: coding agents, single-lineage teachers, and composed engines all beat or match GEPA, and the GEPA team now ships a meta-optimizer on the premise that no single engine wins 3,2,23. The valuable edits are structural: code in Flex, chain restructuring in FAPO, typed protocols in multi-agent systems, the control slot in HARNESSEVO 15,14,18,21. Acceptance is where results are won or lost, and it is still mostly greedy: PACE and PROCTOR state the problem, and STEVE, ESPO, and RLMOpt build partial regression gates into their selection steps 26,1,29,16,24.

The window also clarified where optimized text goes. Anthropic said in July 2026 that it had "removed ~80% of the Claude Code system prompt for our newest models" 33. Instructions a previous model needed became redundant once the next model was trained on the behavior. An optimized prompt's half-life is set by the release cycle, so its value is highest on models that will not be retrained soon (open-weight deployments, frozen enterprise endpoints) and as a source of trajectories and data that the next training run can absorb, which is how FORGE and mmGRPO use it 22,28.

Section 07 / 08

#What ships

The optimizers ship as open-source libraries with humans running the compile. DSPy (about 38,000 GitHub stars in September 2026) includes GEPA, SIMBA, and MIPROv2, and added Flex in August 2026 34,15. GEPA's repository (about 6,700 stars) packages optimize_anything, the omni engines, and the Terrarium benchmarking harness, and claims "50+ production uses" across companies including Shopify, Databricks, Dropbox, and OpenAI 35,3. Vendor features predate the 2026 critiques: OpenAI's dashboard prompt optimizer, Anthropic's Console prompt improver (launched November 2024, per a secondary source), Google's Vertex AI Prompt Optimizer (generally available since August 2025), and Arize's Prompt Learning SDK, built on Andrej Karpathy's May 2025 framing of "system prompt learning" as a missing third learning paradigm 36,37,38,39,40.

The shipped workflow is human-in-the-loop by construction: an engineer assembles data and a metric, runs a compile, reads the result, and deploys. None of these products documents an anytime-valid acceptor, a headroom check, or canary cases. Report 16 (What ships) covers the enterprise improvement loops built on these tools.

Section 08 / 08

#Open problems

Principled acceptance for real systems. PACE gives a guarantee for single-agent prompt edits on small models; PROCTOR gives a production checklist without controlled evidence. Nobody has shown false-commit control for multi-module programs, structural edits, or LLM-judged metrics, which is where FAPO, Flex, and omni operate. Until someone does, the positive results above should be read as upper bounds from greedy gates. CASD sharpens the question: if one unvalidated pass beats validated search, the acceptor has to be justified by the errors it prevents, not assumed.

Transfer across model upgrades. NPO finds prompt transfer within a family, Niklaus finds prose failing across families, and Anthropic's deletions show instructions going stale as models absorb them. No study measures how much of an optimized prompt's gain survives a version upgrade in the same family. The answer decides whether optimized prose is an asset or a recurring cost, and whether lessons belong in code and tests instead, which transfer better.

When to start and when to stop. Coin Flip's headroom test and RLMOpt's finding that headroom, not budget, drives gains give a start rule. omni's plateau-then-switch result suggests a stop-or-switch rule. There is still no convergence theory for reflective optimizers and no standard for how many held-out wins an edit needs before it counts.

Three conclusions survive the caveats. Text optimization works on frozen models when the task has headroom, and it composes with fine-tuning. Candidate generation has become cheap and interchangeable: a coding agent reading the whole trace corpus, or a strong teacher revising one lineage, does as well as elaborate search. A single run on a modern instruction-tuned model is close to a coin flip, graded by an evaluator that a parser fallback or a bad label can fool and an acceptor that commits noise. The work that pays now is checking for headroom before optimizing, committing only on evidence that survives repeated looks, and editing structure when the evidence points there.

Sources

#References

● marks sources dated June 24 to September 24, 2026.

  1. [1]V. Wahi, “LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails,” arXiv 2609.02246, 2026-09-02 (single-author position paper). https://arxiv.org/abs/2609.02246
  2. [2]A. Singh, S. Gautam, P. Gupta, N. Mehrotra, T. Bakshi, S. Gulwani, “Coding Agents are Strong Prompt Optimizers” (CASD), arXiv 2609.26261, 2026-08-13 (preprint). https://arxiv.org/abs/2609.26261
  3. [3]S. Tan, L. A. Agrawal, D. Lee, J. Zhang, D. Klein, K. Sen, A. G. Dimakis, M. Zaharia, “optimize_anything Goes omni: Composing Optimizers into Meta-Optimizer Pipelines,” GEPA blog, 2026-07-22. https://gepa-ai.github.io/gepa/blog/2026/07/22/optimize-anything-omni
  4. [4]Databricks, “Building State-of-the-Art Enterprise Agents 90x Cheaper with Automated Prompt Optimization,” Databricks Blog, 2025-09-24. https://www.databricks.com/blog/building-state-art-enterprise-agents-90x-cheaper-automated-prompt-optimization
  5. [5]Y. Zhou et al., “Large Language Models Are Human-Level Prompt Engineers” (APE), arXiv 2211.01910, 2022-11 (ICLR 2023). https://arxiv.org/abs/2211.01910
  6. [6]C. Yang, X. Wang, et al. (Google DeepMind), “Large Language Models as Optimizers” (OPRO), arXiv 2309.03409, 2023-09 (ICLR 2024). https://arxiv.org/abs/2309.03409
  7. [7]R. Pryzant, D. Iter, J. Li, et al., “Automatic Prompt Optimization with 'Gradient Descent' and Beam Search” (ProTeGi), arXiv 2305.03495, 2023-05 (EMNLP 2023). https://arxiv.org/abs/2305.03495
  8. [8]M. Yuksekgonul, F. Bianchi, et al., “TextGrad: Automatic 'Differentiation' via Text,” arXiv 2406.07496, 2024-06 (Nature Machine Intelligence, 2025). https://arxiv.org/abs/2406.07496
  9. [9]O. Khattab, A. Singhvi, et al., “DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines,” arXiv 2310.03714, 2023-10 (ICLR 2024). https://arxiv.org/abs/2310.03714
  10. [10]K. Opsahl-Ong, M. J. Ryan, et al., “Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs” (MIPROv2), arXiv 2406.11695, 2024-06. https://arxiv.org/abs/2406.11695
  11. [11]L. A. Agrawal, S. Tan, D. Soylu, et al., “GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning,” arXiv 2507.19457, 2025-07 (ICLR 2026 oral). https://arxiv.org/abs/2507.19457
  12. [12]L. A. Agrawal, D. Lee, S. Tan, et al., “optimize_anything: A Universal API for Optimizing any Text Parameter,” arXiv 2605.19633, 2026-05-19 (ACM CAIS 2026). https://arxiv.org/abs/2605.19633
  13. [13]Y. Lee, R. Nair, Q. Zhang, K. Lee, O. Khattab, C. Finn, “Meta-Harness: End-to-End Optimization of Model Harnesses,” arXiv 2603.28052, 2026-03-30. https://arxiv.org/abs/2603.28052
  14. [14]P. Kassianik, B. Saglam, H. Zhao, B. Nelson, S. Vijay, et al., “FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines,” arXiv 2606.19605, 2026-06-17 (preprint). https://arxiv.org/abs/2606.19605
  15. [15]M. Isaac (CMU, at cmpnd), “Introducing Flex: Let the Model Write the Code,” cmpnd blog, 2026-08-05; DSPy docs. https://www.cmpnd.ai/blog/let-the-model-write-the-code.html · https://dspy.ai/current/diving-deeper/flex
  16. [16]L. Liu, P. Tang, K. Y. Singh, S. Ghadar, “ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize,” arXiv 2609.04197, 2026-09-03 (preprint). https://arxiv.org/abs/2609.04197
  17. [17]X. Ma, H. Liu, Y. Li, et al., “SEPO: Evidence-Grounded Prompt Optimization via Structural Editing,” arXiv 2608.28067, 2026-08-28 (preprint). https://arxiv.org/abs/2608.28067
  18. [18]W. Zhang, S. S. Murtaza, J. A. Bhatti, et al., “Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs,” arXiv 2609.00621, 2026-09-01 (preprint). https://arxiv.org/abs/2609.00621
  19. [19]J. Chu, J. Seo, J. Cho, et al., “AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems,” arXiv 2609.08572, 2026-09-08 (preprint). https://arxiv.org/abs/2609.08572
  20. [20]M. Chen, W. Deng, J. Zou, H. Yu, X. Li, “Textual Equilibrium Propagation for Deep Compound AI Systems” (TextEP), arXiv 2601.21064, 2026-01 (preprint). https://arxiv.org/abs/2601.21064
  21. [21]M. Nguyen, W. C. Tan, N. A. Hassan, A. Raman, L. H. Lim, et al., “Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents” (HARNESSEVO), arXiv 2609.02889, 2026-06-25 (preprint). https://arxiv.org/abs/2609.02889
  22. [22]T. Yuan, Z. Qian, “Failure-Guided Co-Evolution of Prompts and Training Data” (FORGE), arXiv 2609.15209, 2026-09-14 (preprint). https://arxiv.org/abs/2609.15209
  23. [23]Y. Chang, X. Chen (Purdue), “Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search,” arXiv 2608.27266, 2026-08-27 (preprint). https://arxiv.org/abs/2608.27266
  24. [24]S. B. Satheesha, N. Pande, D. Duddempudi, B. Dandala, “RLMOpt: Adaptive Prompt Optimization via Recursive Language Models,” arXiv 2608.10471, 2026-08-11 (preprint). https://arxiv.org/abs/2608.10471
  25. [25]P. Singh, “MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization,” arXiv 2607.11944, 2026-07-11 (single-author preprint). https://arxiv.org/abs/2607.11944
  26. [26]Zayx Shawn, “PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents,” arXiv 2606.08106, 2026-06-06 (single-author preprint). https://arxiv.org/abs/2606.08106
  27. [27]X. Zhang, G. Wang, Y. Cui, et al., “Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems,” arXiv 2604.14585, 2026-04-16 (preprint). https://arxiv.org/abs/2604.14585
  28. [28]N. Ziems, D. Soylu, L. A. Agrawal, et al., “mmGRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs,” arXiv 2508.04660, v1 2025-08-06, v2 2026-05-11 (ACM CAIS 2026). https://arxiv.org/abs/2508.04660
  29. [29]Y. Xu, Y. Li, X. Li, et al., “STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification,” arXiv 2609.23716, 2026-09-20 (preprint). https://arxiv.org/abs/2609.23716
  30. [30]J. Niklaus (Hugging Face), “Don't Train the Model, Evolve the Harness,” Hugging Face Space and repo, 2026-07-01. https://joelniklaus-harness-optimization.hf.space/ · https://github.com/JoelNiklaus/harness-optimization
  31. [31]N. Sadjoli, T. Siefken, A. Ghosh, Y. Mai, D. Dahlmeier, “Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading,” arXiv 2604.27637, 2026-04-30. https://arxiv.org/abs/2604.27637
  32. [32]X. Liao, B. Zamanlooy, M. Shafieinejad, et al., “CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization,” arXiv 2607.18622, 2026-07-21 (preprint). https://arxiv.org/abs/2607.18622
  33. [33]T. Shihipar (Anthropic, Claude Code), post on X, 2026-07-24: “We removed ~80% of the Claude Code system prompt for our newest models…” https://x.com/trq212/status/2080710971228918066
  34. [34]Stanford NLP, DSPy GitHub repository (~38,260 stars), accessed 2026-09-24. https://github.com/stanfordnlp/dspy
  35. [35]GEPA project, GitHub repository and README (“50+ production uses”; ~6,730 stars), accessed 2026-09-24. https://github.com/gepa-ai/gepa
  36. [36]OpenAI, “Prompt optimizer,” OpenAI Platform docs, 2025. https://platform.openai.com/docs/guides/prompt-optimizer
  37. [37]Anthropic, “Prompt improver” (Console), 2024-11; launch details via Maginative (secondary). https://www.anthropic.com/news/prompt-improver
  38. [38]Google, “Vertex AI Prompt Optimizer: new algorithm and SDK for general availability,” Google Developer forums, 2025-08-12. https://discuss.google.dev/t/vertex-ai-prompt-optimizer-new-algorithm-sdk-for-the-general-availability/253359
  39. [39]Arize AI, “Prompt Learning” (AX docs) and prompt-learning SDK, 2025-07. https://arize.com/docs/ax/prompts/prompt-optimization · https://github.com/Arize-ai/prompt-learning
  40. [40]A. Karpathy, post on X (“system prompt learning”), 2025-05-10. https://x.com/karpathy/status/1921368644069765486