Research
Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games
Overview Research area: Machine learning, specifically Mean Field Games (MFGs) and Reinforcement Learning (RL) for large-scale multi-agent systems. Technical level: Advanced. The paper assumes familia

- arXiv
- 2602.12517
- Published
- 2026-02-13
- Authors
- Lorenzo Magnino, Jiacheng Shen, Matthieu Geist, Olivier Pietquin, Mathieu Laurière
AI summary
Overview
- Research area: Machine learning, specifically Mean Field Games (MFGs) and Reinforcement Learning (RL) for large-scale multi-agent systems.
- Technical level: Advanced. The paper assumes familiarity with Markov decision processes, best-response maps, Nash equilibria, exploitability, and fixed-point/contraction arguments.
- Scope: The paper proposes Bench-MFG, a standardized benchmark suite of prototypical and randomly generated stationary MFG environments, together with an empirical comparison of learning and optimization algorithms and a set of experimental guidelines.
What This Paper Is About
Work on learning algorithms for Mean Field Games is fragmented: researchers evaluate methods on bespoke, isolated, and often simplistic environments, which makes it hard to compare robustness, generalization, and failure modes. This paper builds a shared evaluation suite for discrete-time, discrete-space, stationary MFGs, organized by mathematical problem class, and adds a generator of random instances so that solvers can be stress-tested statistically rather than on a handful of fixed examples.
Key Contributions
- Taxonomy and prototypical environments. The authors categorize MFGs into No-Interaction, Contractive, Lasry-Lions Monotone, Potential, and Dynamics-Coupled classes, and provide open-source prototypical environments for each (Move Forward, Coordination Game, Beach Bar Problem, Two Beach Bars, Four Room Exploration, Rock Paper Scissor, SIS Epidemic, Kinetic Congestion). The JAX implementation is reported to give 2000x speedup compared to classical open-source methods.
- MF-Garnets procedural generation. A method that constructs random MFG instances with controllable structure and interaction types (additive versus multiplicative dynamics and reward coupling), enabling statistical testing of solvers beyond fixed examples.
- Systematic evaluation. A benchmark of algorithms spanning best-response-based fixed point methods (Fixed Point, Damped Fixed Point, Fictitious Play), policy-evaluation-based methods (Policy Iteration, Smoothed PI, Boltzmann PI, Online Mirror Descent), and exploitability minimization, including a new black-box method, MF-PSO.
- Guidelines. A synthesized set of best practices for reproducible and rigorous MFG experiments.
Main Findings
- Fictitious Play is surprisingly robust. It converges even in non-monotone settings where theoretical guarantees are absent.
- OMD often reaches the lowest exploitability, particularly in cyclic games, but it is highly sensitive to hyperparameters and requires delicate tuning.
- MF-PSO outperforms Fictitious Play in the reported comparisons, at the cost of higher computation that scales linearly with the number of particles.
- Different algorithms behave very differently across game structures and dimensionalities. The MF-Garnet table shows the ranking shifting between the 5x5x5 and 25x10x10 configurations and between dynamics/reward structure combinations. For example, PSO achieves the best mean exploitability in the 5x5x5 (A/M) case (0.2250 ± 0.2115) and in the 5x5x5 (M/M) case (0.1844 ± 0.2449), while OMD is the weakest in the 25x10x10 cases (1.4371 ± 0.3936 for (A/A) and 1.4728 ± 0.4549 for (M/A)).
- Simple fixed point methods can be strong baselines. In the 25x10x10 (A/A) setting, Fixed Point reaches 9.50e-04 ± 0.0017 and Damped FP reaches 8.51e-04 ± 0.0017, competitive with or better than several more elaborate methods.
- Contractive MFGs are rare in discrete spaces. The paper states it is impossible to have interesting MFGs that are contractive when the spaces are discrete; the provided contractive example satisfies the condition at the expense of having a trivial best-response map that is constant with respect to the mean field. Ensuring continuity of the best-response map would require adding an entropy regularization term, which the authors avoid because it changes the Nash equilibrium and departs from classical MFGs.
- Cyclic and non-potential games are distinguishable by structure. Using a linear population reward with interaction matrix A, the game admits a potential if and only if the Jacobian of the population-dependent term is symmetric; skew-symmetric A (as in Rock Paper Scissor) produces non-zero curl and rotational dynamics.
- Multiplicity of equilibria occurs in practice. For the Two Beach Bars problem with parameters satisfying α >> c2 >> c1, the model admits exactly two stationary mean field Nash equilibria, with the population concentrated at one of the two bars.
Methodology in Plain English
The authors fix a single formalism: a finite state and action space, a transition kernel and reward that may both depend on the population distribution, and a discount factor, with a Nash equilibrium defined as a policy whose induced stationary distribution justifies that same policy. They measure how far a policy is from equilibrium using exploitability, the maximum gain an agent could obtain by deviating.
They then sort MFGs by mathematical structure, since structure governs which algorithms can be expected to work. No-interaction games reduce to ordinary MDPs. Contractive games have a unique equilibrium reachable by simple iteration, but the authors note that interesting discrete examples of this class are hard to construct. Lasry-Lions monotone games penalize crowding and satisfy a monotonicity inequality. Potential games admit a potential function and are characterized by a symmetric Jacobian, which lets the authors deliberately build non-potential and cyclic games. Dynamics-coupled games are those where the population changes the transition probabilities, which the authors flag as relatively under-studied and harder because the player's MDP itself shifts during learning.
Beyond hand-designed environments, MF-Garnets generate random instances by sampling a base transition tensor and reward table, plus random coupling coefficients and scalar parameters, then combining base and population-dependent terms in either an additive or multiplicative form, followed by normalization. Because random functions of the mean field cannot be represented with finitely many parameters in general, the authors restrict to these canonical choices.
For the comparison itself, each algorithm is run across the environment suite and over the MF-Garnet instances, with the best hyperparameter configuration selected by average final exploitability across 4 seeds. Experiments use JAX, with a pure Python implementation also provided; all experiments were run on a MacBook Air with an Apple M2 chip (2022).
Why This Matters
Research impact. The paper gives the MFG learning community a shared testbed analogous to what the Arcade Learning Environment, MuJoCo, BenchMARL, and Gorsane et al. (2022) provide for general RL. It also supplies a falsifiable structural taxonomy: an algorithm validated only on monotone congestion games may behave very differently in cyclic or dynamics-coupled games, and the suite makes that visible.
Real-world applications (domains cited in the paper as MFG application areas):
- Finance, including large-population models of trading and market behavior.
- Economics, including population-scale decision problems.
- Epidemiology, including the SIS Epidemic environment where social activity trades off against infection risk that depends on prevalence.
- Crowd and congestion management, including the Kinetic Congestion environment where high population density physically blocks movement.
- Social choice and coordination, including the Two Beach Bars problem inspired by a social choice problem.
Industry relevance. Teams building decentralized or population-scale systems need an efficient way to test whether a solver converges, and JAX-based implementations with JIT compilation and automatic differentiation matter when running exhaustive hyperparameter sweeps. The reported 2000x speedup over classical open-source implementations lowers the cost of statistically meaningful evaluation.
Future Directions
- Extending the suite to large-scale environments to rigorously test algorithm scalability, which the paper explicitly names as the first planned direction.
- Developing contractive or otherwise well-behaved discrete MFG benchmarks whose best-response map is non-trivial, since the current contractive example has a constant best-response map and entropy regularization is ruled out for changing the equilibrium.
- Improving exploitability minimizers such as MF-PSO, whose computational cost scales linearly with the number of particles while delivering better performance than Fictitious Play.
- Broadening the MF-Garnet family of random interaction structures, since the paper notes it restricts itself to a few canonical additive and multiplicative forms and that many variants could be considered.
- The limitations section, as provided, is truncated after the scalability direction, so the remaining planned directions are not reported.
Target Audience
Researchers and graduate students working on Mean Field Games, multi-agent reinforcement learning, and game-theoretic learning who need a standard evaluation protocol. It is also relevant to practitioners implementing or comparing MFG solvers, and to reviewers and benchmark designers interested in how to construct problem taxonomies and random instance generators for algorithmic evaluation.
Authors’ abstract
The intersection of Mean Field Games (MFGs) and Reinforcement Learning (RL) has fostered a growing family of algorithms designed to solve large-scale multi-agent systems. However, the field currently lacks a standardized evaluation protocol, forcing researchers to rely on bespoke, isolated, and often simplistic environments. This fragmentation makes it difficult to assess the robustness, generalization, and failure modes of emerging methods. To address this gap, we propose a comprehensive benchmark suite for MFGs (Bench-MFG), focusing on the discrete-time, discrete-space, stationary setting for the sake of clarity. We introduce a taxonomy of problem classes, ranging from no-interaction and monotone games to potential and dynamics-coupled games, and provide prototypical environments for each. Furthermore, we propose MF-Garnets, a method for generating random MFG instances to facilitate rigorous statistical testing. We benchmark a variety of learning algorithms across these environments, including a novel black-box approach (MF-PSO) for exploitability minimization. Based on our extensive empirical results, we propose guidelines to standardize future experimental comparisons. Code available at \href{https://github.com/lorenzomagnino/Bench-MFG}{https://github.com/lorenzomagnino/Bench-MFG}.