Research
Explaining Decentralized Multi-Agent Reinforcement Learning Policies
Overview Research area: Explainable AI / Explainable Reinforcement Learning, specifically post-hoc explanation of multi-agent reinforcement learning (MARL) policies in decentralized settings. Technica

- arXiv
- 2511.10409
- Published
- 2025-11-13
- Authors
- Kayla Boggess, Sarit Kraus, Lu Feng
AI summary
Overview
Research area: Explainable AI / Explainable Reinforcement Learning, specifically post-hoc explanation of multi-agent reinforcement learning (MARL) policies in decentralized settings.
Technical level: Advanced. The paper assumes familiarity with MARL training paradigms (CTDE and DTDE), partial orders and Hasse diagrams, and Boolean function minimization.
Scope: The paper proposes the first methods for summarizing and answering user queries about decentralized MARL policies, where each agent holds its own policy, acts asynchronously, and observes only its local state.
What This Paper Is About
Existing explanation methods for multi-agent reinforcement learning assume a centralized joint policy executed with full observability, so they cannot handle the uncertainty, nondeterminism, and limited observability of decentralized execution. This paper builds summaries and query answers for decentralized policies by extracting partial-order structure over task completions from agent trajectories, and explicitly marking which dependencies are certain versus merely possible. The goal is to let a human user understand what a team of agents is doing, why they fail to do something, and what they will do next.
Key Contributions
-
Hasse diagram summarization (HDS) of decentralized execution. A novel algorithm constructs, from a single episode of decentralized agent trajectories, a directed acyclic graph whose nodes are sets of simultaneously completed tasks annotated with the agents that performed them, and whose edges encode partial-order constraints over completion times. The paper proves the resulting diagram is both correct (every path projects onto each agent's task sequence consistently) and complete (every agent's full task sequence appears along at least one path).
-
Three types of query-based explanations. Methods for "When do agents G perform task τ?", "Why don't agents G do task τ under conditions Φ?", and "What do the agents do after task τ?" — all designed for decentralized execution.
-
An explicit uncertainty mechanism. Partial comparability graphs identify nodes whose ordering relative to a queried task is unknown, and an uncertainty dictionary records these unordered dependencies. Quine-McCluskey minimization then produces minimal Boolean formulas, translated into natural language where certain conditions are phrased with "must" and uncertain ones with "may."
-
Empirical and human validation. Evaluation across four MARL domains and two algorithms with different training paradigms (centralized training / decentralized training), plus two IRB-approved user studies measuring both objective question-answering accuracy and subjective ratings.
Main Findings
-
Compact summaries versus large baseline visualizations. On the largest configuration of each domain under CTDE policies (SEAC), HDS produced far smaller graphs than the adapted single-agent baseline. Search and Rescue (9 agents, 7 tasks): HDS averaged 8 nodes and 7.88 edges versus 534 nodes and 525 edges for the baseline. Level-Based Foraging (9, 9): 10 nodes and 10.83 edges versus 723 and 714. Multi-Robot Warehouse (4, 19): 20 nodes and 19 edges versus 1,274 and 1,270. Pressure Plate (7, 6): 7 nodes and 6 edges versus 265 and 258.
-
Computational efficiency. Both HDS and the baseline processed 100 episodes and generated summarizations in under one second across all domains. Similarly, both the explanation method and its baseline generated explanations in under one second in all reported cases.
-
Structural regularity behind diversity. SR(9,7) produced 100 unique Hasse diagrams but only 6 distinct edge counts, indicating that the diagrams vary in detail while falling into a small number of structural types.
-
More compact, uncertainty-aware explanations. For "When" queries on CTDE policies, the HD-When method produced far fewer features than the baseline, and it was the only method to report uncertain features. SR(9,7): 9 certain and 2 uncertain features versus 54 certain and 0 uncertain for the baseline. LBF(9,9): 13 and 11 versus 104 and 0. RW(4,19): 0 certain and 153 uncertain versus 267 and 0. PP(7,6): 8 and 3 versus 20 and 0.
-
Uncertainty can dominate. In RW(4,19), all features were marked uncertain because execution is highly asynchronous — the paper notes this is where the baseline, which reports no uncertainty, is least adequate.
-
Summarization study improved accuracy. Users answered significantly more questions correctly with HDS (M = 4.25 out of 6, SD = 0.83) than with the baseline (M = 3.1 out of 6, SD = 1.04); paired t-test t(19) = 4.2, p ≤ 0.01, d = 0.96.
-
Summarization ratings only partially improved. HDS was rated significantly higher only on completeness (Wilcoxon signed-rank W = 16.0, Z = −2.07, p ≤ 0.04, r = −0.33); the other six quality metrics showed no significant difference. Response times were comparable between methods, with HDS sometimes faster (e.g., for likelihood queries).
-
Explanation study improved accuracy on all three query types. "When": t(20) = 9.65, p ≤ 0.01, d = 2.16. "Why Not": t(20) = 13.23, p ≤ 0.01, d = 2.96. "What": t(20) = 12.05, p ≤ 0.01, d = 2.69.
-
Explanation ratings improved on all seven metrics. Understanding (W = 3.5, Z = −3.02, p ≤ 0.01, r = −0.47), satisfaction (W = 5.0, Z = −3.01, p ≤ 0.01, r = −0.46), detail (W = 4.5, Z = −3.02, p ≤ 0.01, r = −0.47), completeness (W = 0.0, Z = −3.21, p ≤ 0.01, r = −0.49), actionability (W = 4.0, Z = −2.53, p ≤ 0.02, r = −0.39), reliability (W = 3.0, Z = −2.78, p ≤ 0.01, r = −0.43), and trust (W = 0.0, Z = −2.68, p ≤ 0.01, r = 0.41). Adding uncertain features did not increase response time.
-
The user studies did not report raw question counts. The paper reports participant counts (20 for summarization, 21 for explanation) and per-question percent-correct style means, but does not report total correct/incorrect answer tallies.
Methodology in Plain English
The researchers treat each agent as running its own policy, acting on its own local observations, with no global clock. As the agents run, they leave behind per-agent trajectories, and the authors extract each agent's task sequence from those trajectories by inferring completed tasks from reward signals and state transitions.
From one episode's worth of these sequences, they build a Hasse diagram: a flowchart-like graph in which each node holds a group of tasks completed at the same time, labeled with the agents that did them, and arrows encode "this happened before that." If two agents complete the same task, they appear in the same node — that is how cooperation is represented. If the ordering between two tasks is not fixed across episodes, the graph simply leaves them unconnected, and different paths through the graph represent the different orderings that actually occur.
To build the diagram, the algorithm walks each agent's task sequence in order, creates nodes for new tasks, adds edges from the previous task to the current one, and finally applies a transitive reduction to remove edges that are implied by longer paths. The worst-case runtime is O(N·|T|² + |T|⁴) for N agents and |T| tasks. A proof shows the result is correct and complete.
For explanations, the authors collect Hasse diagrams from many simulated episodes. To answer a "When" query, they locate the node where the queried task is completed, build a partial comparability graph containing only nodes with a known ordering relative to it, and put the rest into an uncertainty dictionary. Nodes are labeled as targets (the query holds) or non-targets, encoded as Boolean feature vectors, and fed to the Quine-McCluskey algorithm to find the minimal formula distinguishing the two. That formula is rendered into English with a template, using "must" for certain conditions and "may" for uncertain ones. "Why Not" queries swap the target and non-target sets so the formula isolates missing conditions instead. "What" queries look at a node's immediate children for certain successors and at unordered nodes for possible successors, again separating certain from uncertain outcomes.
They tested on four gridworld benchmarks — Search and Rescue, Level-Based Foraging, Multi-Robot Warehouse, and Pressure Plate — where agents see only nearby cells (up to four cells per direction in Pressure Plate, one per direction elsewhere). Policies were trained with SEAC (centralized training) and IA2C (decentralized training), each until convergence or up to 400 million steps, on a 2.1 GHz Intel CPU with 132 GB RAM running Ubuntu 22.04.
Why This Matters
Impact on research. Prior explainable-RL work on multi-agent systems assumed a single centralized controller or independent non-cooperative agents. This paper is the first to summarize and explain agent cooperation and task ordering when execution is decentralized, and it is the first to make ordering uncertainty a first-class, explicitly communicated part of an explanation rather than something hidden behind a definite-sounding answer. It also shows the method is algorithm-agnostic, working on policies trained by both CTDE and DTDE paradigms.
Real-world applications:
-
Search and rescue. A field operator working with a decentralized robot team could ask when the robots will complete a task, why they are not completing it under current conditions, and what they will do next, in order to reprioritize urgent tasks or reallocate resources.
-
Multi-robot warehousing. Warehouse fleets that pick up and deliver items often operate independently for scalability; summaries of task order and cooperation help supervisors diagnose stalls.
-
Autonomous driving. Explaining why a decentralized vehicle policy did not yield or merge can support transparency and safety review.
-
Environments with communication or scalability constraints, where fully centralized control is unavailable and each agent only sees part of the world — precisely the settings where prior centralized explanation methods break down.
Industry relevance. The compactness of the output matters operationally: on the largest tested settings the summaries shrank from hundreds of nodes and edges to roughly 6–20 nodes, and both summary and explanation generation completed in under a second. That makes it plausible to run explanations on live trajectory data during deployment rather than only in post-hoc analysis. The explicit "may" conditions also give system designers a defensible way to communicate what the system genuinely does not know.
Future Directions
-
Interactive human-agent integration. The authors list embedding these explanations directly into interactive human-agent systems, rather than delivering them as static survey outputs, as future work.
-
More expressive query types. The current framework covers "When," "Why Not," and "What"; the authors want to support a broader range of user queries.
-
Large language model integration. The paper proposes leveraging LLMs to improve the clarity and usability of the generated explanations, which are currently produced from fixed structured templates.
-
Open questions the paper leaves. The summarization study found only partial subjective improvement despite better accuracy — the authors attribute this to user familiarity with the baseline's flowchart layout — so how to make partial-order summaries feel as approachable as familiar visual formats remains unresolved. Similarly, when nearly all features are uncertain (as in RW(4,19), where 153 features were uncertain and 0 certain), it is unclear how useful an explanation that is almost entirely "may" conditions can be.
Target Audience
Researchers and graduate students in explainable AI, reinforcement learning, and multi-agent systems will get the most from this paper, particularly those working on post-hoc policy explanation and human-agent teaming. It is also relevant to practitioners deploying multi-robot or multi-agent systems under communication or scalability constraints who need interpretable behavior summaries without access to a central controller. The user-study portions will additionally interest HCI researchers studying how uncertainty should be communicated in AI explanations. Readers should come in with background in MARL training paradigms and basic order theory, since the methods rely on partial orders and Boolean minimization.
Authors’ abstract
Multi-Agent Reinforcement Learning (MARL) has gained significant interest in recent years, enabling sequential decision-making across multiple agents in various domains. However, most existing explanation methods focus on centralized MARL, failing to address the uncertainty and nondeterminism inherent in decentralized settings. We propose methods to generate policy summarizations that capture task ordering and agent cooperation in decentralized MARL policies, along with query-based explanations for When, Why Not, and What types of user queries about specific agent behaviors. We evaluate our approach across four MARL domains and two decentralized MARL algorithms, demonstrating its generalizability and computational efficiency. User studies show that our summarizations and explanations significantly improve user question-answering performance and enhance subjective ratings on metrics such as understanding and satisfaction.