Research
ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment
Overview Research area: AI alignment for large language model (LLM) agents, specifically interpretable reward modeling and multi-agent coordination. Technical level: Advanced. The paper assumes famili

- arXiv
- 2512.06196
- Published
- 2025-12-05
- Authors
- Charlie Masters, Marta Grześkiewicz, Stefano V. Albrecht
AI summary
Overview
Research area: AI alignment for large language model (LLM) agents, specifically interpretable reward modeling and multi-agent coordination.
Technical level: Advanced. The paper assumes familiarity with reinforcement learning from human feedback, policy optimization (GRPO/GSPO), utility theory, and multi-agent systems.
One-sentence scope: ARCANE is a multi-agent framework in which a "manager" agent learns to generate natural-language rubrics that steer "worker" agents toward a stakeholder's preferences, trained with supervised fine-tuning followed by a regularized Group-Sequence Policy Optimization procedure and evaluated on 219 GDPVal-derived tasks.
What This Paper Is About
As LLM-based agents take on longer, project-scale tasks, keeping them aligned with what stakeholders actually want becomes harder, because the reward models used to steer them are usually opaque and fixed at training time. The authors' goal is to make alignment both interpretable (so people can inspect and audit the objectives) and configurable at inference time (so preferences can shift without retraining). ARCANE addresses this by having a manager agent elicit preferences through dialogue and convert them into weighted, verifiable natural-language rubrics that worker agents then follow.
Key Contributions
-
ARCANE framework (Adaptive Rubric-based Control of Agents via Natural-language Exchange): a rubric-based alignment framework in which a manager agent learns to generate rubrics that align worker agents with stakeholder utilities through interactive preference elicitation, framed as a multi-agent collaboration problem.
-
A two-stage learning procedure for the manager policy: initial supervised fine-tuning on synthetic manager–stakeholder dialogues, followed by reinforcement fine-tuning via regularized Group Sequence Policy Optimization (GSPO), with penalty terms for rubric complexity and stakeholder interaction cost.
-
A formalization of alignment as bilevel optimization under partial observability of the stakeholder's latent utility, including a "utility gap" loss and a functional-alignment condition requiring the proxy utility to preserve the stakeholder's preference ordering.
-
Empirical evaluation on GDPVal tasks, showing that learned rubrics (1) guide workers to maximize stakeholder preferences at test time, (2) preserve rankings over outputs consistent with stakeholder preferences, and (3) are interpretable and effectively verifiable.
Main Findings
-
Guided generation beats unguided: Across 44 held-out tasks, mean true return at N=1 rose from 0.58 (No Rubric) to 0.62 (GSPO), with the Oracle Rubric reaching 0.70.
-
The ordering holds at larger sampling budgets: At N=8, mean returns were 0.58 (No Rubric), 0.68 (SFT), 0.74 (GSPO), and 0.81 (Gold/Oracle).
-
Reinforcement fine-tuning beats supervised fine-tuning alone: A one-sided Wilcoxon signed-rank test on the evaluation tasks confirmed statistical significance of GSPO's mean returns over SFT (p=0.0182 at N=8), with mean per-task improvement of +0.044 (W+=592, W−=269; 22 episodes where GSPO wins vs. 19 where SFT wins). At best-of-eight sampling, mean returns were 0.6817 (SFT) vs. 0.7460 (GSPO), significant at α=0.05.
-
Learning curves scale similarly to the oracle: All guided models (SFT, GSPO, and Oracle) exhibit nearly identical best-of-N scaling slopes, with each doubling of N yielding an average ≈ +0.03 absolute gain.
-
The unguided baseline does not improve with more samples: The No Rubric (LLM-Judged) baseline stayed at 0.58 ± 0.01 at N=1 and 0.58 ± 0.02 at N=8, while SFT went from 0.59 ± 0.09 to 0.68 ± 0.03 and GSPO from 0.62 ± 0.12 to 0.74 ± 0.03.
-
Rubrics are compact and legible: Each gold rubric targets 9–12 criteria distributed by weight across three stages (Gate 20–30%, Verification 40–50%, Quality 20–30%), giving an average of roughly 10 criteria per task.
-
Hybrid verification keeps evaluation practical: The rubric corpus contains 2,601 criteria — 1,931 LLM judges (74.2%) and 670 code rules (25.8%).
-
Configurable trade-offs without retraining: Because rubrics expose criteria and weights, stakeholders can directly inspect, edit, or override them at inference time, enabling trade-offs such as correctness versus conciseness.
Methodology in Plain English
The setup has three roles. A stakeholder holds a hidden preference function over outputs. A manager can talk to the stakeholder and learns how to represent those preferences. Workers actually produce the outputs. Because managers often cannot modify worker parameters (workers may be closed-source models or APIs), alignment has to happen through communication rather than parameter sharing.
A rubric is a set of weighted natural-language criteria, each paired with a verifier. A verifier is either a rule-based check (like citation counts or formatting) or a model-based judge (a lightweight LLM or classifier for semantics such as factuality or tone). The rubric's score for an output is the weighted sum of verifier scores, which becomes the manager's proxy for the stakeholder's true utility.
The manager runs a pre-execution consultation: it asks the stakeholder a short sequence of clarifying questions, receives answers expressing priorities and constraints, and then synthesizes a rubric proposal. This interaction is treated as a one-shot cooperative game under partial observability — the stakeholder reveals limited, noisy information through language, and the manager must infer a faithful structured approximation. Because asking questions and running verifiers both cost something, the objective penalizes clarification burden and compute cost alongside reward.
Training proceeds in two stages. First, supervised fine-tuning teaches the decomposer to produce coherent dialogue and rubric structure, using a large reasoning model to synthesize stakeholder responses that yield reference rubrics; loss is masked on system prompts and task inputs. Second, GSPO training treats the manager as a stochastic policy that proposes rubrics. For each task, K rubrics are sampled, each conditions a worker that produces an output, and the stakeholder utility of that output is the scalar return. Advantages are computed by standardizing returns within the group of K samples, and the policy is updated with a length-normalized, sequence-level importance ratio plus a KL penalty keeping it near a reference policy. To improve sample efficiency, the authors add prioritized experience replay at the rubric-group level, replaying episodes whose mean returns fall in the bottom percentile.
At test time, no gradients are needed. Rubric scores rank candidate outputs, enabling best-of-K sampling, importance reweighting, and tree or beam search using rubric scores as heuristics.
For experiments, the authors used the GDPVal corpus, keeping 219 tasks after filtering out unsupported file types (audio, PSD, CAD, Apple Pages) and splitting them into 175 train and 44 evaluation episodes. Tasks require file-based artifacts (.md, .xlsx/.csv, .pdf/.docx, .png/.jpg) so verification can be deterministic. Four baselines were compared, differing only in the source of rubric guidance: Best-of-N with no rubric, SFT-generated rubrics, GSPO-generated rubrics, and an Oracle Rubric given direct access to the ground truth.
Why This Matters
Impact on research. The paper reframes rubric generation itself as a policy-optimization problem rather than treating rubrics as given. This connects interpretable reward modeling to multi-agent coordination, and the authors argue that multi-agent alignment is a property of the interaction process rather than of individual policies. The bilevel formulation and the utility-gap loss give a formal target that connects alignment fidelity to ordinal preference preservation.
Real-world applications (from the paper's task domains):
- Data analysis, where agents manipulate and analyze structured spreadsheets and produce verifiable results.
- Document authoring, where agents generate reports and written artifacts that must follow schema and quality expectations.
- Visual reporting, where outputs such as figures and images must be checked by vision-language judges.
- Code and reasoning tasks, where rubric scores can serve as decoding heuristics for structured search.
Industry relevance. Because the framework requires no worker retraining and works through prompts and shared context, it fits deployments where workers are closed-source foundation models or APIs. Rubrics give organizations an auditable artifact they can inspect and adjust when priorities change, and the compute and clarification cost terms make the trade-off between alignment quality and operational expense explicit.
Future Directions
-
Handling preference drift over time. The paper notes that the true utility may be incompletely specified, revealed only through sparse feedback, or change as preferences drift; how the learned rubrics track such shifts is not resolved.
-
Broadening validation. Results are reported on 44 held-out GDPVal tasks with 219 tasks total after filtering, using a stakeholder simulator for U*; behavior with real stakeholders and other domains remains open.
-
Reducing verifier error. The authors mitigate verifier error by using verifier models from different families than the judged workers and by designing easily verifiable rules, but error in model-based judges is still a limiting factor worth further study.
-
Multi-agent negotiation of criteria. The framework models a one-shot consultation between one stakeholder and one manager; extending this to multiple stakeholders with conflicting preferences, or to negotiated criteria across interacting agents, is a natural next step.
Target Audience
This paper is most valuable to AI alignment and safety researchers, reinforcement learning practitioners working on reward modeling and policy optimization, and engineers building LLM agent systems who need auditable, test-time-adjustable control over agent behavior. Readers will need a working understanding of RLHF, policy gradient methods, and utility theory to follow the methodology sections; the motivation and findings are accessible to a broader technical audience.
Authors’ abstract
As agents based on large language models are increasingly deployed to long-horizon tasks, maintaining their alignment with stakeholder preferences becomes critical. Effective alignment in such settings requires reward models that are interpretable so that stakeholders can understand and audit model objectives. Moreover, reward models must be capable of steering agents at interaction time, allowing preference shifts to be incorporated without retraining. We introduce ARCANE, a framework that frames alignment as a multi-agent collaboration problem that dynamically represents stakeholder preferences as natural-language rubrics: weighted sets of verifiable criteria that can be generated on-the-fly from task context. Inspired by utility theory, we formulate rubric learning as a reconstruction problem and apply a regularized Group-Sequence Policy Optimization (GSPO) procedure that balances interpretability, faithfulness, and computational efficiency. Using a corpus of 219 labeled rubrics derived from the GDPVal benchmark, we evaluate ARCANE on challenging tasks requiring multi-step reasoning and tool use. The learned rubrics produce compact, legible evaluations and enable configurable trade-offs (e.g., correctness vs. conciseness) without retraining. Our results show that rubric-based reward models offer a promising path toward interpretable, test-time adaptive alignment for complex, long-horizon AI systems.