Skip to content
AI.info

Research

Altruism and Fair Objective in Mixed-Motive Markov games

Altruism and Fair Objective in Mixed-Motive Markov Games Overview Research area: Multi-agent systems, specifically multi-agent reinforcement learning (MARL), game theory, and fairness in cooperative d

Altruism and Fair Objective in Mixed-Motive Markov games
arXiv
2602.08389
Published
2026-02-09
Authors
Yao-hua Franck Xu, Tayeb Lemlouma, Arnaud Braud, Jean-Marie Bonnin

AI summary

Altruism and Fair Objective in Mixed-Motive Markov Games

Overview

Research area: Multi-agent systems, specifically multi-agent reinforcement learning (MARL), game theory, and fairness in cooperative decision-making.

Technical level: Advanced. The paper assumes familiarity with normal-form game theory, social welfare functions, Markov games, Bellman operators, contraction mappings, and policy gradient theorems.

Scope: The paper proposes replacing the standard utilitarian welfare objective in cooperative multi-agent learning with a Proportional Fairness objective, defines a "fair altruistic" utility and a Fair Altruistic Markov Game, and derives fair Actor-Critic policy gradient algorithms with experiments planned in the CleanUp environment.

What This Paper Is About

In social dilemmas, agents face a tension between individual self-interest and collective good: each agent is tempted to defect and benefit from others' cooperation without paying the cost, producing unfair outcomes. The dominant approach to multi-agent cooperation — utilitarian welfare — maximizes total reward but can produce efficient yet highly inequitable distributions, the way a utilitarian calculus might justify one person donating a kidney because the recipient gains far more.

This paper asks whether cooperation can be made both efficient and fair by grounding the altruistic objective in Proportional Fairness (the maximization of summed log-utilities) rather than summed raw utilities, and whether that objective can be extended from one-shot normal-form games to sequential, infinitely-horizon Markov games with tractable policy gradient algorithms.

Key Contributions

  1. A fair altruistic utility defined on the log-payoff space. The authors redefine the α-altruistic game extension so that an increasing transformation F_i (they set F_i = log for Proportional Fairness) is applied to each player's shifted payoff, and derive the analytical altruism level α_G needed for cooperation to become a Nash equilibrium in classic social dilemmas.

  2. Analytical cooperation conditions for Prisoner's Dilemma, Stag Hunt, and Chicken. Theorem 3.2 gives α_G = 0 when T ≤ R, and α_G = (log T − log R) / (log R − log S) when T > R, where R is mutual-cooperation reward, T the temptation to defect, and S the sucker's payoff.

  3. A Fair Altruistic Markov Game with infinite horizon. The Proportional Fair state value function V^{π,Prop}(s) = Σ log V_j^{π}(s) is defined over infinite trajectories, using the Ionescu-Tulcea Theorem to construct a consistent probability measure over the infinite trajectory set.

  4. Fair Actor-Critic algorithms via new policy gradient theorems. The paper proves a Fair Policy Gradient Theorem and a Fair Altruistic Advantage Policy Gradient theorem, giving an advantage function A^F_i(s, a, s') = Σ_j c_i(j) A_j(s,a) / V_j(s') with weights c_i(j) = 1 when j = i and α otherwise.

Main Findings

  • Cooperation thresholds are ratios of log-payoffs. The required altruism level in games where cooperation is not initially stable is (log T − log R) / (log R − log S). The numerator, log(T/R), measures the temptation to defect; the denominator reflects the consequences of defection. The authors describe this as the precise degree to which players must internalize the externalities of their actions.

  • Stag Hunt cooperation is never destabilized. Because Stag Hunt has T ≤ R, the ratio (log T − log R) / (log R − log S) is negative while α is always non-negative, so the condition holds for any α in [0, 1]. The transformation never changes the cooperative equilibrium.

  • Prisoner's Dilemma and Chicken both require positive altruism. Both games have T > R, so mutual cooperation (C, C) becomes a Nash equilibrium only once α exceeds the log-payoff ratio. For the Prisoner's Dilemma the preference order is T > R > P > S; for Chicken it is T > R > S ≥ P.

  • Consistency with α ≤ 1 requires TS ≤ R². The authors impose this condition to keep the altruism level within the range consistent with the definition of altruism, noting it holds for standard payoff matrices.

  • The α = 1 case makes all agents share one objective. When α = 1, every agent maximizes the same objective J_i(θ) = E[V^{θ,Prop}(s_0)], and the gradient uses a single shared fair advantage A^F(s, a, s') = Σ_j A_j(s,a) / V_j(s').

  • Experiments are described but not reported in the available content. The paper states that the effects of the fair objective on agent performance and group fairness under various altruism levels are studied in the CleanUp environment (introduced by Hughes et al., 2018 and part of Melting Pot 2.0, Agapiou et al., 2023). The text provided ends mid-sentence at the beginning of the environment description, so no experimental results, numeric outcomes, or baselines are reported here.

Methodology in Plain English

The authors start from an existing idea: take a selfish game and blend each player's own payoff with the group's social welfare, weighted by an altruism parameter α. This is the α-altruistic game. Their complaint is that you cannot simply add raw individual payoffs to a fairness function built from logs, because the two operate on different scales.

Their fix is to push individual payoffs through an increasing function F_i (they choose the logarithm for Proportional Fairness) before mixing, and to shift payoffs so they stay strictly positive — a requirement of the logarithm. With that reformulation, they solve the algebra for the minimum α that makes mutual cooperation a Nash equilibrium, which yields the log-ratio formula. They then apply the same logic to sequences of decisions by treating the expected discounted value of a policy as the "payoff," applying the proportional fairness function to the sum of log-values across agents, and deriving the gradient of this new objective with respect to each agent's policy parameters.

The gradient derivation proceeds by differentiating the Bellman equation, recognizing a contraction mapping, invoking the Banach-Picard fixed-point theorem to guarantee a unique fixed point, and unrolling the recursion. A baseline term is subtracted to reduce variance, producing an advantage-function version of the gradient, from which Actor-Critic updates can be built.

Why This Matters

This work matters because most cooperative MARL research optimizes total welfare and treats fairness as an afterthought applied through environment-specific reward shaping. Reward-shaping approaches tend to reward instantaneous fair behavior and ignore fairness across the full trajectory; related prior work (Ju et al., 2023; Mandal and Gan, 2023) is limited to finite trajectories and lacks a deep policy gradient analysis. This paper instead changes the objective itself, at the level of the log-payoff and log-value space, and provides gradient theorems for infinite-horizon settings.

Real-world applications:

  • Telecommunications and network resource allocation. Proportional Fairness originated in rate control and scheduling; the authors are affiliated with Orange Labs, suggesting direct relevance to spectrum and bandwidth sharing among users or operators.
  • Multi-operator infrastructure sharing. Heterogeneous parties contributing to and drawing from shared network infrastructure face exactly the defection-versus-contribution tension modeled here.
  • Traffic and autonomous mobility coordination. Cooperative routing and intersection management among self-interested vehicles is a mixed-motive sequential problem with fairness concerns among participants.
  • Federated or shared AI systems. Multiple parties training or contributing to a shared model face the question of how benefits and costs are distributed, which is the kidney-donation problem in machine-learning form.

Industry relevance: Any setting where multiple self-interested entities must cooperate repeatedly — cloud resource sharing, collaborative logistics, federated learning, and network slicing — benefits from an objective that balances efficiency with equity instead of maximizing aggregate throughput at the expense of one party.

Future Directions

  1. Actually run and report the CleanUp experiments. The paper promises evaluation across social dilemma environments under various altruism levels, comparing measured fairness and efficiency; the results are absent from the content available.

  2. Determine how to set α in practice. The theory identifies the threshold α at which cooperation becomes an equilibrium, but how an agent should choose α when the game parameters are unknown or changing is not addressed.

  3. Extend beyond two-player, symmetric social dilemmas. The analytical results target symmetric two-player normal-form games defined by R, T, S, and P; general n-player and asymmetric settings remain open.

  4. Compare against reward-shaping and role-based fairness baselines. Prior fairness methods in MARL (Jiang and Lu, 2019; Long et al., 2024, and others) are cited as limitations, but the provided content contains no empirical comparison against them.

Target Audience

This paper is aimed at researchers in multi-agent reinforcement learning and algorithmic game theory who work on cooperation, social dilemmas, and fairness objectives, as well as graduate students comfortable with policy gradient theory and Bellman operator analysis. Applied researchers in telecommunications and networked systems — particularly those working on proportional fair resource allocation — will also find the framing relevant, given the authors' institutional affiliations. Readers looking for empirical benchmark results should note that the experimental section is not included in the material summarized here.

Authors’ abstract

Cooperation is fundamental for society's viability, as it enables the emergence of structure within heterogeneous groups that seek collective well-being. However, individuals are inclined to defect in order to benefit from the group's cooperation without contributing the associated costs, thus leading to unfair situations. In game theory, social dilemmas entail this dichotomy between individual interest and collective outcome. The most dominant approach to multi-agent cooperation is the utilitarian welfare which can produce efficient highly inequitable outcomes. This paper proposes a novel framework to foster fairer cooperation by replacing the standard utilitarian objective with Proportional Fairness. We introduce a fair altruistic utility for each agent, defined on the individual log-payoff space and derive the analytical conditions required to ensure cooperation in classic social dilemmas. We then extend this framework to sequential settings by defining a Fair Markov Game and deriving novel fair Actor-Critic algorithms to learn fair policies. Finally, we evaluate our method in various social dilemma environments.

Read the original paper