Skip to content
AI.info

Research

Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

Overview Research area: Multi-agent systems, reinforcement learning, and diffusion language models applied to structured (JSON) synthetic data generation. Published at AAMAS 2026 (25th International C

Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)
arXiv
2601.07152
Published
2026-01-12
Authors
Aja Khanal, Kaushik T. Ranade, Rishabh Agrawal, Kalyan S. Basu, Apurva Narayan

AI summary

Overview

Research area: Multi-agent systems, reinforcement learning, and diffusion language models applied to structured (JSON) synthetic data generation. Published at AAMAS 2026 (25th International Conference on Autonomous Agents and Multiagent Systems, Paphos, Cyprus, May 25–29, 2026); arXiv:2601.07152v1 [cs.MA], 12 Jan 2026, extended version.

Technical level: Advanced. The paper combines an episodic MDP formulation, a prompt-space policy optimized with PPO/REINFORCE principles, and theoretical results (a contraction-mapping theorem and two propositions on KL divergence and Lipschitz reward ordering).

One-sentence scope: The paper proposes Agents of Diffusion (AoD), a parameter-free multi-agent reinforcement learning loop in which two autoregressive LLM agents (a prompt optimizer and a judge) steer a frozen diffusion language model toward JSON outputs that are both schema-conformant and semantically diverse.

What This Paper Is About

Generating structured data such as JSON records is hard for large language models: autoregressive models keep the format but tend to repeat themselves and lose diversity, while diffusion language models produce richer, more varied text but lack the inductive biases needed to preserve nested structure. AoD's goal is to get both properties at once by letting autoregressive LLM agents guide a frozen diffusion generator through natural-language feedback, without fine-tuning the diffusion model, handcrafted constraints, or scalar reward shaping.

Key Contributions

  1. The authors introduce Agents of Diffusion, described as the first multi-agent reinforcement learning framework to guide diffusion language models using natural language.
  2. They propose an optimization loop in which LLM agents iteratively refine prompts through verbal critique, achieving schema-aligned control without reward modeling and without updating the diffusion model's weights.
  3. They report reproducible state-of-the-art results on JSON-based instruction synthesis across four structured datasets (MultiWOZ, Super-NaturalInstructions, TruthfulQA, Self-Instruct), each randomly subsampled before training and evaluation to reduce memorization risk.
  4. They provide theoretical support: Theorem 1 argues the prompt-update sequence forms a contraction mapping in expectation for sufficiently small step size, and propositions on KL divergence (diffusion closer to the real data manifold than autoregressive under bounded reconstruction error) and on the Lipschitz/monotone behavior of the judge-based reward operator.

Main Findings

  • Best task success, least memorization: Across all datasets and prompt-optimizer/judge LLM pairs, AoD reaches the highest Task Success Rate (0.79) and the lowest Field Overlap (0.29) among all compared systems. Each experiment was repeated 15 times.
  • Balanced similarity and diversity: AoD scores 0.88 on Similarity, 0.82 on Diversity, 0.83 on Novelty, and 6.10 on Entropy, with low Perplexity (22.1). The prose discussion elsewhere states Diversity as 0.72, which differs from the 0.82 value listed in Table 1.
  • Strong independent metrics: AoD leads on BLEU (35.6), ROUGE-L (40.1), and METEOR (29.3), which the authors use as independent evaluation metrics rather than training feedback.
  • Comparison to the diffusion backbone alone: The baseline LLaDA (Nie et al., 2025) records Similarity 0.79, Diversity 0.69, Novelty 0.81, Entropy 6.03, Perplexity 27.0, BLEU 29.5, ROUGE-L 34.2, METEOR 25.8, TSR 0.69, and Field Overlap 0.35.
  • Comparison to diffusion baselines: Diffusion-LM (Li et al., 2022) records Similarity 0.72, Diversity 0.60, Novelty 0.72, Entropy 5.82, Perplexity 29.4, BLEU 28.1, ROUGE-L 33.5, METEOR 25.1, TSR 0.61, and Field Overlap 0.42; DiffLM (Zhou et al., 2024) records Similarity 0.74, Diversity 0.63, Novelty 0.70, Entropy 5.90, Perplexity 28.6, BLEU 27.5, ROUGE-L 32.9, METEOR 24.6, TSR 0.63, and Field Overlap 0.41.
  • Static autoregressive baselines fall short on diversity and success: LLaMA-3.1 8B, for example, records Similarity 0.86 but Diversity 0.42, Novelty 0.48, Entropy 5.18, Perplexity 21.6, BLEU 33.9, ROUGE-L 38.1, METEOR 27.9, TSR 0.71, and Field Overlap 0.38.
  • Frozen generator: The diffusion model LLaDA-8B is never updated. Controllability comes only from the autoregressive prompt policy, which the authors describe as parameter-free with respect to the generator and model-agnostic.
  • Judge design matters: The judge cluster combines an LLM judge with a Natural Language Evaluator that computes similarity, Distinct-n, entropy, novelty, and perplexity and converts them into textual statements; the prompt optimizer never sees scalar rewards directly, which the authors argue limits collusion and reward hacking.
  • Hardware accessibility: All experiments ran on a consumer-grade workstation with an AMD Ryzen 9 7900X (12-core, 24-thread, 4.7 GHz base), 32 GB DDR5 RAM, and an NVIDIA RTX 4080 SUPER GPU (16 GB VRAM), alongside API endpoints.

Methodology in Plain English

AoD sets up a loop. A frozen diffusion language model (LLaDA-8B, 32 layers, sinusoidal embeddings, 1024-token input, T = 12 denoising steps, FP16, classifier-free guidance disabled) generates a JSON candidate from a current prompt. A judge agent evaluates that candidate against a rubric using a fixed set of yes/no questions, supported by five quantitative signals computed by a Natural Language Evaluator: semantic similarity to the reference set, Distinct-n diversity, token entropy, novelty, and perplexity. A scorer turns the critique into a scalar reward plus a vector of subrewards.

A prompt-optimization agent, instantiated as an autoregressive LLM, then proposes a natural-language edit to the prompt based on a summary of the interaction history, and the edit operator applies it. This repeats for a bounded number of outer iterations, with the policy updated using a blend of proximal policy optimization and REINFORCE principles. The diffusion weights stay fixed; only prompts change. Eight autoregressive models were tested in the agent roles: LLaMA-3.1 8B (32 layers, 40 heads, 4-bit quantization; temperature 0.7 for prompting and 0.2 for judgment), Qwen-3 8B (nucleus sampling with p = 0.9), DeepSeek-R1 8B (top-k sampling with k = 40), Gemma-2 9B (beam search width 3), Mistral 7B (8-bit decoding), and the API models GPT-4.1 Nano, GPT-4.1 Mini, and GPT-4.1. The same autoregressive model served as both prompt optimizer and judge within each experiment.

Why This Matters

Impact on research. The paper argues that structure control and generative diversity need not be traded off, and that diffusion language models can be supervised through language alone rather than gradient updates or handcrafted grammars. It also reframes prompt optimization from an offline search problem into an online reinforcement learning problem, connecting multi-agent coordination research to the emerging Diffusion Language Models literature.

Real-world applications (as implied by the paper's framing).

  • Synthetic JSON record generation for training or augmenting models when real data is limited, sensitive, or costly.
  • Schema-constrained data synthesis for nested structured formats, such as the booking-style example with departure_city, arrival_city, and date fields.
  • Instruction-tuning data creation, via the JSON instruction synthesis tasks drawn from Super-NaturalInstructions and Self-Instruct.
  • Factuality-sensitive question answering data, represented by the TruthfulQA evaluation used to probe hallucination and semantic precision.

Industry relevance. The authors emphasize that AoD runs on consumer-grade hardware and supports both open-weight local models and proprietary API models, making controllable structured generation feasible without large-scale compute clusters. The released code is at https://github.com/Idsl-group/AgentsOfDiffusion.

Future Directions

  • Extending beyond JSON: The paper works with JSON and nested schemas; whether the same verbal-feedback loop transfers to other constrained symbolic formats is not reported in the provided content.
  • Scaling the agent roles: Experiments held the prompt optimizer and judge to the same autoregressive model for consistency. Whether mixing models or using specialized judges improves results remains an open question.
  • Reward-signal verification: The authors propose Field Overlap and paired monitoring of similarity versus novelty trends as diagnostics for leakage and collusion, but the provided content does not report an explicit external audit of the learned reward landscape.
  • Formal guarantees in practice: The contraction and Lipschitz claims are stated under local smoothness and bounded-edit assumptions; the provided text does not report empirical tests of how tightly those assumptions hold.

No explicit future-work section appears in the truncated content provided.

Target Audience

Researchers and practitioners in multi-agent systems, reinforcement learning, and generative modeling who are interested in controllable synthetic data generation; engineers building structured-data pipelines who want schema-conformant output without fine-tuning; and readers with a background in policy-gradient methods or diffusion models who want to see how the two can be combined through natural-language feedback.

Authors’ abstract

Generating high-quality structured data such as JSON records, remains a fundamental challenge for large language models (LLMs), particularly when semantic richness must coexist with strict schema adherence. While autoregressive LLMs offer strong structural consistency, they often struggle with semantic variation and output diversity. In contrast, diffusion language models (DLMs) introduce powerful mechanisms for semantic richness and bidirectional decoding, yet lack the inductive biases needed for reliable structure preservation. We present Agents of Diffusion (AoD), a novel framework that unifies the generative flexibility of DLMs with the reasoning capabilities of autoregressive models through language-mediated reinforcement learning. AoD frames structured text generation as a multi-agent alignment process, where a prompt optimization agent collaborates with a judge agent to iteratively guide a DLM using natural language feedback. This approach enables controllable, schema-consistent generation without modifying model parameters or relying on handcrafted constraints. AoD advances the state of controllable generation by demonstrating that diffusion models, when supervised by cooperative agents, can achieve both high semantic novelty and structural fidelity. Across multiple structured data benchmarks, AoD consistently outperforms diffusion and autoregressive baselines, establishing a new path forward for structure-aware, diversity-enhanced text synthesis.

Read the original paper