Research
Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
Relic: From Multi-Agent Collaboration to Persistent Organizational Capability Overview Research area: Multi-agent AI systems — specifically multi-agent software engineering, agent coordination, and th

- arXiv
- 2609.32965
- Published
- 2026-09-26
- Authors
- Hongyi Du, Tianyi Zhang, Weijia Zhang, Yi Yang, Haofei Yu, Kunlun Zhu, Tianxiang Dai, Shang Jiang, Zhelun Gao, Jiaxin Pei, Shang Zhu, Jiaxuan You
AI summary
Relic: From Multi-Agent Collaboration to Persistent Organizational CapabilityOverview
Research area: Multi-agent AI systems — specifically multi-agent software engineering, agent coordination, and the persistence of organizational knowledge across changing team membership.
Technical level: Advanced. The paper assumes familiarity with LLM-based multi-agent teams, agent benchmarks, and evaluation protocols for software tasks.
Scope: The paper proposes Relic, a mechanism that converts recurring multi-agent collaboration failures into organization-owned, executable protocols, and evaluates whether those protocols improve delivery and survive the departure of the agents who authored them.
What This Paper Is About
In multi-agent software teams, agents often work against each other: one agent changes an interface while another keeps building on the old version, leaving tests stale and integration broken. A conversation among the agents can resolve that particular episode, but once the participants change, nothing guarantees the lesson carries forward to the next team. The paper's goal is to make those lessons a durable property of the organization rather than of the individual agents, so that later work is governed by rules the team itself authored and adopted.
Key Contributions
- A protocol lifecycle for multi-agent teams. Relic is introduced as a system in which members reflect on visible work, propose rules, and participate in governing whether those rules are adopted — turning recurring collaboration failures into organization-owned artifacts.
- Executable protocols rather than advisory text. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences directly to the runtime, and remain open to revision and retirement as work continues.
- An evaluation across controlled runs and a public benchmark. The work reports 360 controlled runs over ten software workloads and three models, plus results on the CooperBench benchmark and a fresh-member transfer experiment.
- Evidence that protocols transfer to new members. The paper isolates the value of executable bindings relative to the same rules supplied as readable text, testing whether inherited rules still govern behavior when the original authors are gone.
Main Findings
- Protocols raise complete-contract delivery in controlled runs. Across 360 controlled runs spanning ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76%, a gain of 5.71 percentage points over a matched structured team that lacks the protocol lifecycle. The abstract states that all four verified production endpoints improved in every model stratum, but does not describe what those endpoints are.
- Execution bindings outperform text alone under member turnover. In a fresh-member transfer setting, behavioral correctness is 25.4% with no inherited protocol, 34.6% when the same rules are supplied as readable text, and 41.2% with executable bindings — a 6.5-point advantage over text alone.
- Strong results on CooperBench. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic reaches 367/477 (76.9%), which the authors describe as the best reported result among peer-structured systems.
- Coordination loss is reversed on a same-model subset. On the fixed 48-pair same-model subset, Relic scores 29/48 versus Solo's 26/48, reversing the coordination loss exhibited by the official peer baseline.
- A traced case shows the lifecycle in action. In one traced case, repeated integration friction produced an interface-review rule that went on to govern later pull requests and was revised as work continued.
Methodology in Plain English
The approach is organizational rather than purely architectural. Agents first inspect work that is visible to the team. When a collaboration failure recurs — for example, repeated integration friction around an interface — members reflect on it and propose a rule. The team then governs whether that rule is adopted. An adopted protocol is not just documentation: it is bound into the runtime, so it specifies what triggers the rule, who is responsible, what evidence is required, and what happens at execution time. Protocols are not frozen; they can be revised or retired as work proceeds.
To test the idea, the authors ran 360 controlled comparisons across ten software workloads and three models, using a matched structured team without the protocol lifecycle as the comparison point. They also ran a transfer experiment in which new members inherit a protocol in one of three forms — nothing, readable text, or executable bindings — to separate the value of writing rules down from the value of making them enforceable. Finally, they evaluated on the CooperBench benchmark and on a fixed 48-pair same-model subset against Solo and the official peer baseline.
Why This Matters
Impact on research. Most multi-agent work focuses on how agents communicate within a single episode or task. This paper shifts attention to what persists after the episode ends and after the participants are replaced, framing collaboration experience as organizational state. That reframing introduces a measurable distinction between rules that are merely readable and rules that are executable, and it provides a benchmark-grounded argument that peer-structured multi-agent systems need not lose to solo agents.
Real-world applications:
- Software engineering teams with high agent or contributor turnover. Interface-review rules that survive personnel changes reduce repeated integration breakage.
- Long-running automated pipelines. Recurring failures at handoff points can be encoded as gated checks with named responsibilities and required evidence instead of being re-litigated each cycle.
- Organizations deploying multiple vendor or model agents together. Executable protocols give a model-agnostic layer of governance when the underlying agents differ.
- Compliance and audit contexts. Because protocols record triggers, responsibilities, and required evidence, they produce a traceable account of why a rule exists and when it was revised or retired.
Industry relevance. The paper's central claim — that executable bindings transfer better than text, and that inherited protocols improve correctness for new members — speaks directly to enterprises that want accumulated operational knowledge to outlive the individual agents, prompts, or people that produced it.
Future Directions
- How protocols should be revised and retired over time. The abstract notes that protocols remain open to revision and retirement but does not describe the mechanisms or criteria for either.
- What determines which failures become protocols. Identifying recurring failure patterns worth codifying, and avoiding a proliferation of low-value rules, is left open.
- Generalization beyond the studied settings. The results cover ten software workloads, three models, CooperBench, and a 48-pair same-model subset; behavior in other domains, larger organizations, or heterogeneous model mixes is untested here.
- Separating the lifecycle from the runtime binding. The transfer experiment isolates executable bindings from text, but the contributions of reflection, proposal, and governance stages relative to each other are not decomposed in the abstract.
Target Audience
Researchers and practitioners working on multi-agent LLM systems, agent coordination, and agentic software engineering; engineers designing long-lived agent organizations that need durable governance; and evaluation-focused researchers interested in benchmarks for persistent organizational capability rather than single-episode task success.
Authors’ abstract
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%), establishing the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.