Skip to content
AI.info

Research

It Takes Workflows to Evolve Better Workflows

Overview Research area: Natural Language Processing — multi-agent LLM systems, workflow generation, and reinforcement learning. Technical level: Advanced. The paper assumes familiarity with reinforcem

It Takes Workflows to Evolve Better Workflows
arXiv
2610.01026
Published
2026-10-01
Authors
Xuehang Guo, Haoyu Wang, Haifeng Chen, Yangyi Chen, Zhenhailong Wang, Qingyun Wang

AI summary

Overview

Research area: Natural Language Processing — multi-agent LLM systems, workflow generation, and reinforcement learning.

Technical level: Advanced. The paper assumes familiarity with reinforcement learning (policy gradients, PPO-style clipped surrogates), credit assignment, and multi-agent orchestration.

Scope: The paper proposes FloWright, a training and test-time paradigm that uses a single shared "workflow harness" to optimize one or several of the roles that build and execute LLM workflows, plus DataWright, a data-hardening method that turns single-agent datasets into workflow-level tasks.

What This Paper Is About

Complex real-world tasks can exceed what a single large language model (LLM) can do, which motivates multi-agent workflows that coordinate specialized agents. Recent methods train LLMs to build better workflows from execution outcomes, but they optimize only the workflow generator while the agents that build or execute each workflow stay fixed, even though every outcome depends on all of them. The paper's goal is to extend training beyond the generator despite two obstacles: agents are coupled with one another, and a workflow's outcome is a single sparse score that cannot say which agent caused a failure. A second goal is that workflows are commonly trained and evaluated on data a single agent can already handle.

Key Contributions

  1. A shared workflow harness for every role, at train and test time. The paper defines one shared harness whose signal serves three purposes: the evaluation metric, the optimization signal for training any role, and the meta distillation objective at test time.
  2. Structure-aware credit at no extra cost. A reward hierarchy converts the sparse signal of a workflow into dense, role-level feedback, with credit read from the execution trace the harness already records — requiring no additional models, labels, or executions.
  3. Workflow-level data hardening (DataWright). Three hardening strategies convert datasets a single agent can already handle into workflow-level tasks with increased difficulty levels, supporting both training and evaluation.
  4. Four evolution modes that generalize. FloWright enables (a) single-role self-evolution, (b) agent-skill co-evolution, (c) upstream-downstream co-evolution, and (d) multi-agent co-evolution; gains are reported as generalizable across roles, backbones, RL algorithms, workflow topologies, and pools.

Main Findings

  • Untrained FloWright already beats single-agent baselines by at least +11.39% overall. Smaller models building and running workflows outperform the same-size model acting alone, including baselines both without and with tool and skill pools.
  • Every evolution mode improves over its untrained counterpart, by up to +7.41%. At matched model size, each mode exceeds the untrained FloWright in every regime and overall, and every arm improves under at least one mode.
  • Co-evolving more roles gains more than optimizing one alone. Multi-agent co-evolution reaches +5.03% overall, ahead of agent-skill co-evolution at +4.06%, upstream-downstream co-evolution at +3.06%, and single-role self-evolution at +2.83%. Multi-agent co-evolution gains +7.32% in-distribution, +6.05% out-of-distribution, and +4.20% out-of-domain, and leads 7/8 columns of the main results table among the four Qwen3.5-4B modes.
  • Test-time distillation also helps, and compounds with training. Distilling a reusable prior from the system's own experiences gains +0.93% overall; from a stronger teacher (GPT-5.4), +1.10%; meta distillation gives the largest test-time gain at five shots with up to +2.58%, while one shot stays below the untrained baseline. Combining test-time and train-time optimization reaches +3.21% overall, +4.14% in-distribution, +3.80% out-of-distribution, and +2.79% out-of-domain, exceeding the five-shot prior alone by +1.59%.
  • Gains transfer across roles and backbones. An optimized 4B role transfers its gain when other roles run on Qwen3.5-9B, even raising accuracy above the all-9B baseline. The meta-distilled prior shows persistent gain in every regime on all four evaluated backbones, including the two it never optimizes (GPT-5-mini and GPT-5.4-mini).
  • The training signal generalizes across RL algorithms. Optimizing the Generator with GRPO gives +2.83% overall, CISPO +3.23%, and DAPO +3.84%; DAPO leads in every regime and widens its margin out-of-domain (+3.69% versus +2.82% for CISPO and +2.59% for GRPO), while staying within 0.01% of CISPO in-distribution.
  • Every layer of the reward ladder contributes, with validity contributing the most. Grading format and answer alone gives only +0.07% under GRPO and +0.80% under DAPO; adding the validity term contributes the largest single increment at +2.07% and +2.29% respectively; the structure-aware credit adds a further +0.69% and +0.75%, reaching +2.83% and +3.84%.
  • Optimization improves every workflow topology. Code as the workflow representation reaches 34.12% untrained and 35.86% with an optimized Generator, ahead of schema, state machine, and graph in both. Optimization improves schema by +2.83%, state machine by +2.62%, code by +1.74%, and graph by +1.25%, narrowing the spread across the four representations from 3.91% to 3.03%.
  • Capsule pools lead at every backbone, but optimization narrows the gap. Bundling each agent with the tools and skills it needs gains +3.83% on Qwen3.5-4B and +3.87% on Qwen3.5-9B over a modular pool; optimizing the Generator improves the modular pool by +4.50% versus +2.83% for capsules, narrowing the gap to +2.16%.
  • A pool that grows on demand gives the largest single design effect. Growing the pool on demand gains +4.83% over a fixed pool. Distilling past experiences into the pool adds a further +1.44% when only successful experiences are kept, against +0.53% when both successful and failed experiences are kept.

Methodology in Plain English

The authors treat a workflow as a directed graph whose nodes are agents, tools, or skills drawn from a pool, and whose edges carry data and control flow. A "harness" runs that graph on a task, returns an answer, and scores it between 0 and 1. That score is used three ways: as the evaluation metric, as the training signal, and as the objective at test time.

To avoid needing a separate reward model or extra runs, the harness reuses the execution trace it already records. Each role — the Generator that designs the workflow, the Inventor that creates missing components, a downstream agent that executes a node — is charged only for the failures of the nodes it authored: the credit is a negative number proportional to the fraction of that role's own nodes that root-cause a failure, and zero if none fail.

The single scalar score is then staged into a hierarchy of automatically verified terms: does the role's output parse, is the role's contribution valid, does the workflow execute to a final answer, how correct is that answer, plus the structure-aware credit term. Weighting coefficients are configured per role. Because one harness serves every role, a single role can self-evolve, or two or more roles can co-evolve, where each role's behavior reshapes the training distribution the others see — a "mutual curriculum."

At train time the paper maximizes the expected signal with policy-gradient methods, anchoring on the PPO clipped surrogate while leaving the advantage estimator and policy loss interchangeable (GRPO, DAPO, and CISPO all instantiate it). At test time no weights change; instead FloWright optimizes a reusable prior distilled internally from the system's own experiences, externally from a stronger teacher, or both, and additionally distills successful experiences into few-shot cases for meta learning.

To supply harder data, DataWright converts 12 datasets over 7 domains into workflow-level tasks using three hardening strategies at hardening levels ℓ ∈ {3, 5}. Each dataset is split from the source before hardening to prevent train-test overlap. Each triple of dataset, hardening strategy, and hardening level is called an "arm," yielding 44 arms in total, grouped into in-distribution, out-of-distribution, and out-of-domain regimes relative to the four training arms (paired strategy, ℓ = 3).

Why This Matters

Impact on research. The paper argues that prior workflow-optimization work trains only the generator while every other agent that builds or executes the workflow remains fixed, leaving much of the workflow's potential improvement on the table. It also argues that credit assignment for multi-agent systems typically costs extra models, labels, or repeated executions — and that existing attributors identify the decisive failure step in only about a third of failed runs — whereas FloWright's role-level credit is read from a trace the harness already produces. Finally, it argues that common evaluation data (single-question math, function-level code, short question answering) can neither reveal nor optimize what a workflow adds beyond a single agent, and it offers DataWright as a replacement.

Real-world applications (drawn from the domains and task types the paper evaluates):

  • Document understanding and processing pipelines over heterogeneous inputs.
  • Slide and chart generation, where subtasks can be parallelized or branched.
  • Code and math problem solving that benefits from decomposition and iteration.
  • Finance tasks, as one of the seven domains covered by the 44 evaluation arms.

Industry relevance. The trained models are small open models (Qwen3.5-4B and Qwen3.5-9B), so the approach targets settings where building and running workflows is cheaper than relying on a much larger single model. The paradigm is optimizer-agnostic and the gains are reported to hold across RL algorithms, workflow topologies, pool granularities, and pool dynamics, which matters for teams already using a particular RL stack. The finding that an optimized 4B role transfers gains to a 9B downstream role suggests mixed-size deployments can be improved without retraining every component.

Future Directions

  • Broadening the model families evaluated. The authors state that their reinforcement learning experiments are mainly evaluated on two open models of a single family and that they aim to incorporate more model families.
  • Extending to dynamic environments. The current 44 arms are tasks whose inputs are fixed once the task begins; the authors plan to extend FloWright to settings where the inputs a workflow reads and the components it needs change while the workflow runs.
  • Training roles at larger scale. The authors list this explicitly as future work.
  • Open question: which role combinations to co-evolve. The paper shows that co-evolving more roles gains more than optimizing one alone, but it does not report a general rule for choosing which roles to pair, and its dynamics ablations show that the benefit of distilling past experience depends on whether successful experiences only or both successful and failed ones are kept.

Target Audience

Researchers and practitioners in multi-agent LLM systems, reinforcement learning for LLM agents, and workflow automation who want to move beyond training a single workflow generator. It is also relevant to evaluation-focused researchers, since DataWright reframes how workflow-level tasks are constructed and scored. Readers with no background in policy-gradient RL will find Sections 3.2 through 3.5 dense, while the results tables and ablations are readable without it.

Authors’ abstract

Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to $+7.41\%$, with co-evolving ($+5.03\%$) more roles gaining more than optimizing one of them alone ($+2.83\%$). Our project page: https://xhguo7.github.io/FloWright/.

Read the original paper