#Six graphs
On September 2, 2026 a UCLA group posted a system that picks a multi-agent topology for each incoming query from a codebook 1. The group used a vector-quantized autoencoder to compress successful agent graphs into discrete codes, then varied how many codes the model was allowed to keep. Growing the codebook from 8 entries to 64 did not buy variety: the topologies that survived training collapsed to about six distinct graphs no matter how much capacity was on offer 1. The final system, Codebook Agent, uses a 16-entry codebook, a reward-weighted MLP that maps a query embedding to a distribution over codes, and a second MLP that reranks decoded candidates. It posts the best accuracy on all six benchmarks in the paper (84.6 average against 83.0 for the strongest prior method), picks a topology in 2.4 ms, and spends 21.9–33.2% fewer LLM tokens than the graph generators it replaces 1.
That result lands on an assumption the subfield had carried since 2024: that the space of multi-agent organizations is large, that the right organization depends on the query, and that finding it justifies a search procedure. If six graphs cover what works, the elaborate conditional graph generators built over the previous two years were spending their expressiveness on distinctions that did not matter. The authors also say why: the message-passing scorers those generators use are invariant to adjacency when every agent shares the same profile, so a scorer cannot tell two wirings apart when every node runs the same model with a similar prompt 1.
September brought two more results pointing the same way. A study of sparse multi-agent debate found that letting each agent debate two freshly sampled distinct peers each round, with no learned topology at all, "provides a surprisingly strong baseline and consistently improves the accuracy-cost trade-off," and concluded that learned topology adaptation should be tested against simple routing and stopping baselines before its complexity is justified 2. An audit of methods that make multi-agent systems cheaper by pruning agents or edges found that many reported gains are setup-dependent and "may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy" 3.
Codebook Agent is a pro-search paper that wins by searching less. Read next to those audits and the budget-controlled studies of early 2026, which find a single agent matching homogeneous multi-agent systems at equal compute 4,5, it supports a narrower conclusion. For a team of copies of one model, the fine structure of the graph carries little information. What does carry information is coarser: how much collaboration a given problem deserves, and which genuinely different models to call. The systems with large, durable gains learned the second with a small trained orchestrator. The production system that most resembles "an agent designs its own team" does no offline search at all.
#Why anyone searches for topologies
A multi-agent system (MAS) is a harness with more than one model call in its control flow: roles (planner, coder, critic), edges (who reads whose output), and an execution order. The hand-designed catalog (debate, ensembles, reflection loops, role-playing software teams) was long by 2024, and choosing among its entries was manual labor. Automated design treats that choice as optimization. The artifact being improved is the workflow, stored outside the weights as code, a graph, or text, so in the series' vocabulary this is non-parametric self-improvement: a proposer suggests a new organization, an evaluator scores it on held-in tasks, and an acceptor keeps the better one.
The lineage is short. ADAS (August 2024) had a meta-agent write new agents as Python code, keep the good ones in an archive, and iterate; its discovered agents beat hand-designed baselines and, unusually, kept their advantage on held-out domains and models 6. AFlow (October 2024) replaced open-ended generation with Monte Carlo Tree Search over code-represented workflows and let small models beat GPT-4o on specific tasks at 4.55% of the cost 7. MaAS (February 2025) stopped searching for one design and learned a distribution over architectures, sampling a cheap system for easy queries and a larger one for hard queries at 6–45% of the inference cost of prior systems 8. By mid-2025, per-query design was the default, whether by an inference-time design-critique loop (MAS-Zero 9) or by a trained generator that emits a system in one pass. Everything in 2026 builds on that framing.
The naive version fails for three reasons that shaped what came after. The search space is combinatorial, since every role, prompt, edge, and loop bound multiplies the options. Evaluation is expensive, since scoring one candidate means running it over a validation set with several LLM calls per example. And the acceptor is weak: most papers keep the design that scores highest on a validation split drawn from the same benchmark they report on, which is the same greedy gate the series keeps finding in prompt optimization and harness evolution (see 06, Prompt and program optimization, and 09, Harness self-improvement). A search that runs long enough against a small validation set will find organizations that fit it. The field answered the first two problems with better representations and cheaper search. It mostly deferred the third, and 2026 is when the bill arrived.
#How topologies get chosen in 2026
Graph generators keep multiplying
The graph-generation line is still the busiest. In August 2026 RGA-Designer took an existing autoregressive topology generator, trained a reward model that scores both task correctness and structural compactness, and fine-tuned the generator against it RLHF-style; it matched the original's accuracy while cutting token consumption by 20.5% on average 10. K-GAT (August 2026) conditions graph generation on retrieved evidence rather than the model's parametric knowledge alone, and reports beating an LLM-Debate baseline on GPQA by 15.7% in accuracy while using less than half the tokens 11. MAGIC (September 2026) builds a graph step by step, choosing for each functional role whether to instantiate a single agent or a reusable group, and trains the construction policy with RL plus potential-based reward shaping to cope with sparse final-answer rewards; it reports beating state-of-the-art baselines across eight benchmarks 12.
The pattern across these papers is telling. Their headline improvements are mostly cost reductions at equal accuracy, and their baselines are other graph generators or fixed debate setups, not a single agent given the same tokens. Codebook Agent's collapse explains why cost is where the wins show up: if the useful set of graphs is small, a better generator mostly learns to stop adding edges.
Selecting how much to collaborate
The sharpest 2026 paper on what topology choice buys takes the coarse view. In September 2026 Yunsong Hong evaluated five code-generation topologies, from a single agent to hierarchical collaboration, on 614 problems from APPS, HumanEval+, and LiveCodeBench 13. The advantage of hierarchical collaboration over a single agent grew from 2.4 points of pass@1 on the easiest third of problems to 21.1 points on the hardest third, while its token cost stayed about ten times higher throughout. A math study on 400 problems reproduced the shape, 2.5 points widening to 20.9 13.
The proposed selector, DATS (Difficulty-Aware Topology Selector), predicts each topology's chance of solving a problem and picks the one that maximizes predicted success minus a cost penalty. Because the penalty is one scalar, every router can be recalibrated to the same spend, and under that budget-matched protocol six cost-aware methods spread across 21.6 percentage points; two baselines that led DATS fell behind once calibrated to its budget 13. Fixed at 40% of the always-hierarchical cost, DATS reached 77.7% pass@1 against 73.6% for always-hierarchical and 74.3% for the strongest learned competitor, with the 4.1-point gain holding across four backbones. Replacing its 39 interpretable features with a graph network or a pretrained encoder moved accuracy by at most 1.3 points 13.
DATS gives the budget critique its most useful refinement. Collaboration is not worthless under a matched budget; it is worth a great deal on hard problems and almost nothing on easy ones, and the router that knows which is which is simple. The decision that matters is one-dimensional (how much structure to spend), which fits a world where six graphs cover the useful space.
Building the graph at inference time
A second line constructs or edits the workflow during execution. MANTA (July 2026) initializes a task-conditioned topology, monitors collaboration traces, and makes bounded structural edits (roles, links, order, visibility, validation paths) while holding the agent budget fixed, reaching a 74.0 average across five benchmarks, 5.8 points over the strongest baseline 14. ReActNet (September 2026) compiles a query into a sequence of directed communication graphs, one per reasoning stage, where every edge carries a natural-language instruction for what the source agent should tell the target; it is training-free and reports improving over fixed-topology and learned-topology baselines at competitive cost 15. Earlier in the year AOrchestra (February 2026) abstracted every sub-agent as a tuple of instruction, context, tools, and model and instantiated one per subtask, reporting a 16.28% relative gain on GAIA, SWE-Bench, and Terminal-Bench with Gemini-3-Flash 16. None of the three reports a matched-compute single-agent comparison; MANTA's fixed agent budget is the closest concession.
Writing design knowledge as text
A third line stores what search learns as prose. ABSTRAL (March 2026) keeps the MAS architecture in an evolving natural-language document refined by contrastive analysis of successful and failed traces; on SOPBench (134 bank tasks with a deterministic oracle) it reached 70% on validation and 65.96% on test with GPT-4o, and a transferred design document matched a cold start's third iteration after one iteration in a new domain 17. Skill-MAS (June 2026) stores orchestration capability as an evolvable "Meta-Skill" and reports, without headline numbers in its abstract, that Meta-Skills transfer to unseen tasks and LLMs 18. Both apply the series' finding that lessons written as inspectable text are cheap to keep and move (see 08, Skills and tools).
#Learning the orchestrator instead of the graph
The systems with the largest gains stopped treating the workflow as the object of search. They train a small model to act as a manager: at each step it decides which worker to call, with what instruction, and which prior messages that worker sees. The workers are frozen models, often from different labs. The policy lives in the orchestrator's weights and what it writes is text for other models, so these are hybrids (report 12, Consolidation and co-evolution, owns that framing; they appear here because they are where topology search went).
Fugu: the orchestrator as a product that keeps retraining
Sakana's Fugu is the clearest case, and its summer 2026 releases show the loop that keeps such a system current. Fugu is built on Conductor, a 7B model RL-trained to write natural-language workflows over a pool of frontier workers, choosing which agent to call, the subtask it gets, and which earlier messages it can see; training randomizes the worker pool, and the Conductor may call itself, which yields recursive orchestration 19,20. The ICLR 2026 paper reported 83.9% on LiveCodeBench and 87.5% on GPQA-Diamond, above every individual worker in its pool, from a modest RL budget of 200 iterations of 4 questions with 64 rollouts each 19. Fugu went generally available on June 22, 2026 as a single model API that internally calls a pool of LLMs, itself included, with training that combines fine-tuning, evolutionary algorithms, and RL 21,22.
On July 24, 2026 Sakana released Fugu-Ultra v1.1, "upgraded to incorporate the latest frontier models," with gains of up to 7.9 points over v1.0 at the same price, and added Claude Code-compatible endpoints so the orchestrated pool can serve as the model behind Claude Code 23. On September 11, 2026 came two more: Fugu Max, which routes to "the leanest model capable of solving" a task across an expanded pool of open and specialized models including NVIDIA Nemotron, priced at $2 per million input and $6 per million output tokens (output pricing Sakana puts 40–60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3), and Fugu Ultra v2, which Sakana says reaches the best or joint-best score on five of eight benchmarks, including 48.3 on the Chartography visual-reasoning benchmark against 27.3 for Opus 5 and 29.5 for Fable 5 24. Sakana stresses that Fable 5, Fable 5.1, and GPT-6-Astra are not in Ultra v2's pool (training cutoff August 28, 2026). All of these numbers are vendor-reported, several on Sakana's internal benchmark, and none is independently replicated.
The release cadence is the finding. The orchestrator's policy encodes the strengths of a specific worker pool, so Sakana re-trains or re-tunes it as the pool changes: three model versions in under three months, each described as an orchestration upgrade over new workers. That is the open problem of pool drift answered, for now, by a vendor running the retraining loop itself.
Research orchestrators: cost-aware routing over heterogeneous executors
The research versions make the cost term explicit. EASy (August 2026) trains an LLM orchestrator with RL and gives it explicit capability and cost profiles for a heterogeneous set of executors; it decomposes tasks into milestones, builds dependency-aware execution graphs, assigns executors, parallelizes independent steps, and is trained with tree-structured rollouts and rewards for correctness, efficiency, and trajectory completeness, reporting better performance-efficiency trade-offs than strong agentic baselines on math, embodied, and deep-research benchmarks 25. The precedent is NVIDIA's ToolOrchestra (November 2025), whose 8B orchestrator scored 37.1% on Humanity's Last Exam against GPT-5's 35.1% at 2.5× lower cost 26. Kimi K2.5 (February 2026) brought the recipe inside a frontier training run: its Parallel-Agent RL trains an orchestrator to spawn and assign parallel sub-agents while keeping them frozen, so credit flows only to the orchestrator; the swarm scores 78.4% on BrowseComp against 60.6% for single-agent K2.5 and runs 3–4.5× faster on WideSearch as target item-F1 rises from 30% to 70% 27 (report 01, Agentic RL, covers the training).
Orchestration is also becoming something models are measured on. SwarmBench (August 2026) evaluates models as swarm orchestrators on accuracy, efficiency, cost, and process quality, finds substantial differences in orchestration capability across current models, and shows that a simple experience-extraction and replay method improves orchestration consistently 28. A capability that differs by model and improves with replayed experience is a capability labs can train, which is where PARL already went.
#What the evidence shows
Sorted by what they compare against, the results fall into bands, and the band tells you how much to trust the headline.
| Date | System | Mechanism | Headline result | Baseline and budget context | Ref |
|---|---|---|---|---|---|
| Sep 2026 | Distinct-peer random routing | No learned topology in sparse debate | Consistently better accuracy-cost trade-off than prior sparse debate | Learned and fixed debate topologies | 2 |
| Sep 2026 | Topology vs diversity study | 2×3 controlled matrix, fixed 3-call budget | Agent differentiation shifts results more than topology; best 52.83 Macro-F1 (Qwen2.5-14B) | Same backbone, same call budget | 29 |
| Sep 2026 | DATS | Difficulty-aware selection among 5 topologies | 77.7% vs 73.6% pass@1 at 40% of always-hierarchical cost; collaboration gain 2.4 → 21.1 pts easy → hard | Budget-matched protocol; 4 backbones | 13 |
| Sep 2026 | Codebook Agent | 16-entry topology codebook | 84.6 vs 83.0 avg; 2.4 ms; 21.9–33.2% fewer tokens | Graph generators; homogeneous agents | 1 |
| Sep 2026 | Fugu Ultra v2 / Fugu Max | Trained orchestrator over model pool (hybrid) | Ultra v2: Chartography 48.3 vs 27.3 (Opus 5); Max: best overall on 6 benchmarks | Vendor-reported; frontier models absent from Ultra v2 pool | 24 |
| Aug 2026 | K-GAT | Evidence-conditioned graph generation | +15.7% accuracy over LLM-Debate on GPQA at < half the tokens | Debate baseline | 11 |
| Aug 2026 | RGA-Designer | Reward-model fine-tuned graph generator | Same accuracy as base generator, −20.5% tokens | Prior generator | 10 |
| Jul 2026 | Fugu-Ultra v1.1 | Orchestrator re-tuned for new workers | Up to +7.9 points over v1.0 | Vendor-reported, same price | 23 |
| Jul 2026 | MANTA | Inference-time topology edits | 74.0 avg, +5.8 pts over strongest baseline, 5 benchmarks | Automated MAS baselines; fixed agent budget | 14 |
| Apr 2026 | SAS vs MAS | Matched reasoning-token study | Single agent matches or beats MAS on multi-hop reasoning | 3 model families, equal thinking tokens | 5 |
| Feb 2026 | Kimi K2.5 PARL | RL orchestrator, frozen sub-agents | BrowseComp 78.4% vs 60.6% single-agent; 3–4.5× faster on WideSearch | Same model, single-agent mode | 27 |
| Jan 2026 | Rethinking MAS | Budget-controlled study | Single agent matches homogeneous MAS on 7 benchmarks | Same model, KV-cache reuse | 4 |
| Dec 2025 | Conductor | RL-trained 7B orchestrator (hybrid) | LiveCodeBench 83.9%, GPQA-Diamond 87.5%; beats every worker | Individual workers; no matched-sample single-worker baseline | 19 |
| Nov 2025 | ToolOrchestra | RL-trained 8B orchestrator (hybrid) | HLE 37.1% vs GPT-5 35.1% at 2.5× lower cost | Single frontier model; cost reported | 26 |
The first band compares a generated topology against other topologies. Most 2026 graph papers sit here (K-GAT, RGA-Designer, MAGIC, MANTA, Codebook Agent), and their gains are real in the narrow sense that the new design beats the old one on that benchmark, mostly by spending fewer tokens. The comparison cannot tell you whether the organization mattered or whether the baseline was wasting calls.
The second band compares against strong single models with cost reported: Conductor, ToolOrchestra, and the Fugu releases. These are the most consequential results because they beat the alternative a practitioner would reach for. All of them route among different models, so their advantage plausibly comes from model selection rather than graph structure, and the Fugu numbers are vendor-reported.
The third band holds compute fixed. The January and April critiques find the homogeneous multi-agent advantage disappears. DATS adds the positive half: at matched spend, collaboration pays on hard problems and a cheap difficulty router captures that. The September topology-versus-diversity study adds the other half: under a fixed three-call budget on multilingual emotion detection, how agents are differentiated (sampling, role prompts, or separately fine-tuned specialists) moved results more than how they were connected 29. Searching the wiring of identical agents buys efficiency; deciding how much to collaborate and who the collaborators are buys capability.
#Where it breaks
The single-agent baseline catches up
The budget critique arrived in two steps in early 2026. Xu et al., in an ICLR 2026 paper, ran homogeneous workflows (one base model playing several roles) against a single agent across seven benchmarks in coding, math, QA, domain reasoning, and planning; the single agent matched them, with an efficiency advantage because one context reuses its KV cache where separate agents re-encode shared information 4. The authors named their own boundary: a single model cannot simulate a team of genuinely different base models. In April 2026 Dat Tran, Douwe Kiela, and colleagues supplied the theory 5. By the data processing inequality, passing information through extra processing steps cannot increase what it tells you about the answer, so under a fixed reasoning-token budget and perfect context use, a single agent is at least as information-efficient as a multi-agent arrangement of the same model. Across Qwen3, DeepSeek-R1-Distill-Llama, and Gemini 2.5 on multi-hop reasoning, single agents matched or beat multi-agent systems at equal thinking tokens; multi-agent systems became competitive only when context use degraded or they got more compute 5.
The strongest reply is that "perfect context use" fails on exactly the tasks where teams help. Salesforce's MASBENCH (January 2026) varies tasks along depth, horizon, breadth, parallelism, and robustness and finds multi-agent gains depend on task structure rather than holding universally 30; Kimi's swarm results on wide search 27 and DATS's hard-problem gains 13 are cases where one context stops being enough. The critiques do not show teams are useless. They show that a homogeneous team on a single-context task buys nothing a single agent with the same budget could not, and most topology benchmarks are that setting.
Efficiency gains that are really collapse
The September efficiency audit sharpens the point for the cost-reduction papers 3. It re-ran representative pruning and compact-structure methods under a shared backbone, agent registry, and runtime, across controlled variations in topology, scale, depth, and tool use, on tasks chosen to demand a multi-agent system. Many reported gains turned out to depend on method-specific prompts and starting topologies, and some came from the method collapsing the structure, disabling tool pathways, or starting from a system so redundant that random pruning preserved accuracy too. The random-routing debate study reaches the same conclusion from the other side: a policy with no learned structure is a strong baseline that learned topology control often fails to beat 2. Together they are topology search's version of report 09's July 2026 finding that evolved harnesses did not consistently beat simple test-time scaling once someone ran the matched baseline 31 (report 14, Measuring self-improvement, covers the protocols).
Handoffs lose the rules, not the facts
Teams also fail at the edges between agents. MAST (March 2025) found in 1,600+ traces that multi-agent failures cluster into design issues, inter-agent misalignment, and verification failures 32. An August 2026 study located one mechanism precisely 33. When an upstream interaction is compressed into a handoff summary, the summary keeps operational facts and drops the boundary metadata that governs how those facts may be used, a failure the authors call "summary collapse." Uncompressed free-text handoffs preserved boundary markers at a survival score of about 0.80; a 25-word budget dropped it to about 0.57 while fact survival stayed near ceiling, and the two were nearly uncorrelated. Vague constraints leaked in 73% of GPT-5-mini cases and 50% of DeepSeek-R1-32B cases, and explicit constraints cut leakage below 15% across the three models tested. A single-agent control still leaked, so the failure is not purely topological, but every added edge is another compression step. Topology search optimizes task accuracy and cost; nothing in its objective sees what a handoff dropped.
Search overfits the benchmark it searches on
Almost every system in the lineage searches on a validation split of a benchmark and reports on the test split of the same benchmark, and the RL orchestrators train and evaluate on the same benchmark families 19. ADAS's cross-domain transfer 6 and ABSTRAL's design-document transfer 17 are exceptions. ABSTRAL also measured the cost of coordination directly: its ensemble configurations achieved only 26% turn efficiency 17. A searched topology that adds agents pays that tax on every query; a codebook of small graphs pays less, which is one reading of why Codebook Agent's token savings and accuracy arrived together 1.
#State of play, 2026
The last three months split the subfield three ways. Graph generation kept producing papers (RGA-Designer, K-GAT, MAGIC, Codebook Agent), with gains that are mostly token savings against other generators. Evaluation papers arrived to check them: the September efficiency audit, the random-routing baseline, the topology-versus-diversity matrix, and DATS's budget-matched protocol. And orchestration moved into products, with three Fugu releases in three months and Claude Code tuning its workflow defaults in September. The weight of evidence moved from "find the right graph" to "decide how much structure to spend, and on which models."
Two earlier-2026 results frame what the window did not settle. CORAL (April 2026) ran several long-lived agents asynchronously on open-ended optimization problems in isolated workspaces with a shared memory of attempts, notes, and skills, every attempt a git commit scored by a separate grader; it reported state-of-the-art results on ten tasks with 3–10× higher improvement rates than fixed evolutionary search, including a kernel benchmark cut from 1,363 to 1,103 cycles by four agents 34. Its structure is flat and hand-designed; its gains come from the grader and the shared memory, which puts it closer to evolutionary search with parallel workers (see 11, LLM-guided evolutionary search) than to org-chart design. And Hetero-Designer (ACL 2026) framed the automated design of teams built from different LLMs as its own problem 35, the regime the critiques leave open, where no budget-matched study yet exists.
#What ships
Production systems do not run offline topology search. The closest thing in a shipped product is Claude Code's dynamic workflows, released with Claude Opus 4.8 in late May 2026 36,37. When a task is large enough, Claude writes a JavaScript harness for it using primitives such as agent(), pipeline(), and parallel(); the harness decomposes the job into phases, fans work out to isolated sub-agents, and synchronizes results. A run executes up to 16 agents concurrently by default, up to 1,000 agents in total, and up to 4,096 items per fan-out 36. Anthropic's announcement is titled "a harness for every task" 37, and the phrasing is accurate: the topology is generated fresh by the model for each job, not retrieved from a searched library.
The product kept moving through September. Version 2.1.269 (September 11, 2026) added a setting to raise the per-run concurrency limit anywhere from 1 to 256 agents, and version 2.1.271 (September 14) changed the default workflow size to small on Pro plans and lowered the "medium" size guideline from 15 to 10 agents 38. The current documentation describes an ultracode effort level under which Claude plans a workflow for every substantive task in a session, bundled workflows such as /deep-research whose verifier agents cross-check claims, and distribution of saved workflows through plugins 36. The September changes point one way: smaller default teams, with scale available when a task needs it, which is the DATS lesson arrived at by product tuning.
Loop closure is configurable, not fixed. Launching a workflow is an ordinary permission-gated tool call. Default manual and accept-edits modes prompt the user before each run. Bypass-permissions mode, headless claude -p runs, and Agent SDK configurations can remove that prompt; auto mode lets a classifier approve it 36. A model can therefore design and launch a thousand-agent job with no human click if the operator has configured it that way. The human-curated part lives elsewhere: when a run goes well, the user can save it with a keystroke as a reusable command under .claude/workflows/, building a library of vetted harnesses 36. The proposer is the model; the acceptor for the library is a person.
Fugu, covered above, is the clearest commercial case found for this series of a learned policy deciding the org chart per request, and its July Claude Code-compatible endpoints put the two products in the same terminal 23. Its trajectory confirms the direction of the evidence: what shipped is a trained router over heterogeneous models, retrained as the pool changes, not a searched graph.
The major agent frameworks are hand-designed. The OpenAI Agents SDK composes agents through explicit handoffs 39, Google's ADK offers sequential, parallel, and loop workflow agents that developers wire by hand 40, and LangGraph defines graphs in code 41. None of the major frameworks runs topology search by default. The gap between research and practice here is less a lag than a verdict: frontier models write a decent harness for a task on the fly, and the budget evidence says an offline-searched graph over copies of one model would rarely beat it.
#Open problems
Budget-matched benchmarks for heterogeneous pools. The critiques cover homogeneous teams; the positive results come from heterogeneous routing. DATS shows what a budget-matched protocol looks like for one model and five topologies 13. Nothing comparable exists for a learned router over several frontier models against the best single model given the same dollars and tokens, sampled as many times, on held-out tasks. Until it does, Conductor-style and Fugu-style results sit in the second evidence band, and nobody can separate how much of their advantage is model selection, how much is extra calls, and how much is structure. The same benchmark would test whether Codebook Agent's six-graph collapse survives once agents differ.
Deployment evidence. Claude Code's dynamic workflows and Fugu are the two large deployments of model-designed organizations, and neither has published success rates, failure taxonomies, or cost per task against single-agent runs. Whether thousand-agent runs launched without a human prompt fail in the ways MAST and the handoff study describe (unverified handoffs, dropped constraints) is an empirical question with safety weight (see 15, Failure modes and safety).
Orchestrators when the model pool changes. A routing policy encodes the strengths of the workers it was trained on. Conductor randomizes pools during training 19, and Sakana's summer releases show a vendor re-tuning the orchestrator every month or two as new workers arrive 23,24. Whether that retraining is cheap enough to keep up, whether routing policies transfer across model generations, and whether an orchestrator can adapt in context from a few trial calls decide whether this hybrid stays economical. It is the orchestrator-specific form of the series question about whether lessons survive a model upgrade.
The reader should come away with a narrower belief than the subfield started with. Searching the wiring of identical agents mostly finds small graphs and saves tokens. Deciding how much to collaborate on a given problem is worth a lot and takes a simple router. Learning whom to call among different models produces the largest gains, and it lives in orchestrator weights that must be retrained as the models change. The systems that ship have the model write a new harness for each task, keeping human curation for the harnesses worth saving.
#References
● marks sources dated June 24 to September 24, 2026.
- [1]J. Yu, Y. Li, et al. (UCLA), “Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems,” arXiv:2609.02264, 2026-09-02. https://arxiv.org/abs/2609.02264
- [2]B. Wang, Z. Li, X. Huang, Y. Dong, “Do We Need Complex Topology Control? Distinct-Peer Random Routing Improves Cost-Efficiency in Sparse Multi-Agent Debate,” arXiv:2609.27150, 2026-09-22. https://arxiv.org/abs/2609.27150
- [3]J. Zhang, L. Zhang, P. Lu, Q. Zhang, Y.-N. Chuang, Z. Li, et al., “Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems,” arXiv:2609.05933, 2026-09-05. https://arxiv.org/abs/2609.05933
- [4]J. Xu et al., “Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline,” ICLR 2026, arXiv:2601.12307, January 2026. https://arxiv.org/abs/2601.12307
- [5]D. Tran, D. Kiela, et al., single-agent vs multi-agent LLMs on multi-hop reasoning under matched reasoning-token budgets, arXiv:2604.02460, April 2026. https://arxiv.org/abs/2604.02460
- [6]S. Hu, C. Lu, J. Clune, “Automated Design of Agentic Systems,” arXiv:2408.08435, August 2024. https://arxiv.org/abs/2408.08435
- [7]J. Zhang et al., “AFlow: Automating Agentic Workflow Generation,” arXiv:2410.10762, October 2024 (v4 April 2025). https://arxiv.org/abs/2410.10762
- [8]G. Zhang, L. Niu, J. Fang, K. Wang, L. Bai, X. Wang, “Multi-agent Architecture Search via Agentic Supernet” (MaAS), arXiv:2502.04180, February 2025. https://arxiv.org/abs/2502.04180
- [9]“MAS-Zero: Designing Multi-Agent Systems with Zero Supervision,” arXiv:2505.14996, May 2025. https://arxiv.org/abs/2505.14996
- [10]P. Suwannapichat, B. Changaival, C. Wu, P. Bouvry, “Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design,” arXiv:2608.20099, 2026-08-20. https://arxiv.org/abs/2608.20099
- [11]Y. Jiang, J. Fan, M. Xu, Y. Guo, J. Feng, S. Xu, et al., “When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems” (K-GAT), arXiv:2608.27984, 2026-08-28. https://arxiv.org/abs/2608.27984
- [12]K. Yang, Z. Yi, X. Li, M. An, Z. Liu, Z. Chen, et al., “MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning,” arXiv:2609.26667, 2026-09-22. https://arxiv.org/abs/2609.26667
- [13]Y. Hong, “Learning How Much to Collaborate: Difficulty-Aware Topology Selection for Multi-Agent Code Generation” (DATS), arXiv:2609.13890, 2026-09-12. https://arxiv.org/abs/2609.13890
- [14]M. Huang, J. Wang, et al. (Cornell), “MANTA: Multi-Agent Network Topology Adaptation,” arXiv:2607.28527, July 2026. https://arxiv.org/abs/2607.28527
- [15]K. Tieu, D. Fu, Y. Xia, H. Li, H. Yan, J. He, “Inference-Time Graph Engineering for Multi-Agent LLM Workflows” (ReActNet), arXiv:2609.05774, 2026-09-04. https://arxiv.org/abs/2609.05774
- [16]J. Ruan et al., “AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration,” arXiv:2602.03786, February 2026. https://arxiv.org/abs/2602.03786
- [17]W. Song, J. Yue, Z. Pang, “ABSTRAL: Automatic Design of Multi-Agent Systems Through Iterative Refinement and Topology Optimization,” arXiv:2603.22791, March 2026. https://arxiv.org/abs/2603.22791
- [18]H. Lin, Q. Yang, C. Qin, “Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems,” arXiv:2606.18837, June 2026. https://arxiv.org/abs/2606.18837
- [19]S. Nielsen, E. Cetin, P. Schwendeman, Q. Sun, J. Xu, Y. Tang (Sakana AI), “Conductor: learning to orchestrate frontier models with RL,” ICLR 2026, arXiv:2512.04388, December 2025 (v5 May 2026). https://arxiv.org/abs/2512.04388
- [20]Sakana AI, “Learning to Orchestrate,” blog, April 2026. https://sakana.ai/learning-to-orchestrate
- [21]Sakana AI, “Fugu” release announcement, 2026-06-22. https://sakana.ai/fugu-release
- [22]Y. Tang, E. Cetin, S. Nielsen, et al. (Sakana AI), “Fugu Technical Report,” arXiv:2606.21228, June 2026. https://arxiv.org/abs/2606.21228
- [23]Sakana AI, “Announcing Fugu-Ultra v1.1 and Claude Code interface for Fugu,” blog, 2026-07-24. https://sakana.ai/fugu-1-1-claude-code-interface/
- [24]Sakana AI, “Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier,” blog, 2026-09-11. https://sakana.ai/fugu-max-release/
- [25]J. Liu, L. Luo, T.-T. Vu, G. Haffari, “EASy: Towards Efficient LLM-Based Agentic System,” arXiv:2608.04588, 2026-08-05. https://arxiv.org/abs/2608.04588
- [26]H. Su, S. Diao, et al. (NVIDIA), “ToolOrchestra” (Orchestrator-8B), arXiv:2511.21689, November 2025. https://arxiv.org/abs/2511.21689
- [27]Kimi Team (Moonshot AI), “Kimi K2.5: Visual Agentic Intelligence,” arXiv:2602.02276, 2026-02-02 (revised 2026-08-07). https://arxiv.org/abs/2602.02276
- [28]J. Gao, Z. Jin, T. Men, K. Liu, J. Zhao, “SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?,” arXiv:2608.30661, 2026-08-31. https://arxiv.org/abs/2608.30661
- [29]U. Shernazarov, C. R. W. Basnayake, A. El Jarjini, N. Crespi, P. Rajapaksha, “Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection,” arXiv:2609.14570, 2026-09-13. https://arxiv.org/abs/2609.14570
- [30]Z. Ke, Y. Ming, et al. (Salesforce Research), “MAS-Orchestra” (with MASBENCH), arXiv:2601.14652, January 2026. https://arxiv.org/abs/2601.14652
- [31]Y. Wang, H. Zhu, Z. Hu, et al., “Rethinking the Evaluation of Harness Evolution for Agents,” arXiv:2607.12227, 2026-07-14 (v2 2026-08-27). https://arxiv.org/abs/2607.12227
- [32]“Why Do Multi-Agent LLM Systems Fail?” (MAST), arXiv:2503.13657, March 2025. https://arxiv.org/abs/2503.13657
- [33]Y. Wang, A. Goyal, E. Chandrasekharan, H. Sundaram, “Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs,” arXiv:2608.29028, 2026-08-29. https://arxiv.org/abs/2608.29028
- [34]MIT CSAIL / CMU / MIT Media Lab, “CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery,” arXiv:2604.01658, April 2026. https://arxiv.org/abs/2604.01658
- [35]Z. Zhang, Y. Zhang, et al. (CASIA), “Hetero-Designer: Automated Design of Multi-Agent Systems with Heterogeneous LLMs,” ACL 2026. https://aclanthology.org/2026.acl-long.1272
- [36]Anthropic, “Workflows,” Claude Code documentation, accessed 2026-09-24. https://code.claude.com/docs/en/workflows
- [37]T. Shihipar (Anthropic), “A harness for every task: dynamic workflows in Claude Code,” Claude blog, 2026-06-02. https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code
- [38]Anthropic, Claude Code changelog, v2.1.269 (2026-09-11) and v2.1.271 (2026-09-14). https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md
- [39]OpenAI, “OpenAI Agents SDK” documentation. https://openai.github.io/openai-agents-python/
- [40]Google, “Agent Development Kit: Multi-agent systems,” documentation. https://google.github.io/adk-docs/agents/multi-agents
- [41]LangChain, “LangGraph.” https://www.langchain.com/langgraph ### Other systems in the lineage (not discussed above) | System (date) | What it adds | Ref | |---|---|---| | GPTSwarm (Feb 2024; ICML 2024) | Agents as optimizable computational graphs; separate optimizers for node prompts and edges | https://arxiv.org/abs/2402.16823 | | EvoAgent (Jun 2024) | Evolutionary operators grow a single-agent framework into multi-agent variants | https://arxiv.org/abs/2406.14228 | | EvoMAC (Oct 2024; ICLR 2025) | “Textual backpropagation” from test failures into collaboration-network structure for software tasks | https://arxiv.org/abs/2410.16946 | | AgentSquare (Oct 2024) | Modular search over planning/reasoning/tool/memory modules; 17.2% avg gain over best human designs, 6 benchmarks | https://arxiv.org/abs/2410.06153 | | ScoreFlow (Feb 2025) | Score-DPO workflow generator; +8.2% over baselines, 6 benchmarks | https://arxiv.org/abs/2502.04306 | | MASS (Feb 2025; ICLR 2026) | Google's staged optimization: block prompts → topology → global prompts | https://arxiv.org/abs/2502.02533 | | MAS-GPT (Mar 2025) | Trained model emits a full MAS as code in one call | https://arxiv.org/abs/2503.03686 | | FlowReasoner (Apr 2025) | RL-trained query-level meta-agent; +10.52% over o1-mini on 3 code benchmarks | https://arxiv.org/abs/2504.15257 | | W4S (Apr 2025) | 7B meta-agent RL-trained in ~1 GPU-hour to design workflows for stronger models | https://arxiv.org/abs/2504.04785 | | Puppeteer (May 2025) | RL-trained orchestrator; compact cyclic structures emerge | https://arxiv.org/abs/2505.19591 | | OFA-MAS (Jan 2026; WWW 2026) | One mixture-of-experts graph generator across tasks | https://arxiv.org/abs/2601.12996 | | EvoMAS (May 2026) | Policy-gradient workflow adapter along a single task trajectory | https://arxiv.org/abs/2605.08769 | | RoboPhD (Apr 2026) | Elo-tournament evolution vs GEPA vs greedy under a 1,500-evaluation budget; ARC-AGI 27.8% → 65.8% | https://arxiv.org/abs/2604.04347 | | OptiMAS (Aug 2026; EMNLP 2026) | Trajectory-driven continuous MAS optimization with dual-track memory | https://arxiv.org/abs/2608.21918 | | EvoAgentX (Jul 2025, v2 Sep 2025) | Open-source platform packaging AFlow, TextGrad, MIPRO for offline workflow optimization | https://arxiv.org/abs/2507.03616 |