Research
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory Overview Research area: AI safety evaluation for multi-agent settings, combining game theory (canonical 2×2 games, Nash equil

- arXiv
- 2602.12316
- Published
- 2026-02-12
- Authors
- Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin
AI summary
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game TheoryOverview
Research area: AI safety evaluation for multi-agent settings, combining game theory (canonical 2×2 games, Nash equilibria, welfare functions) with large language model benchmarking.
Technical level: Intermediate. The paper assumes familiarity with basic game-theoretic concepts (Nash equilibria, best responses, payoff matrices) but explains its welfare metrics and methodology in accessible terms; the appendices carry the formal machinery.
Scope: The paper builds and releases GT-HarmBench, a 1,535-scenario benchmark that maps real-world AI safety risks onto six canonical symmetric 2×2 games and uses it to measure, diagnose and partially mitigate unsafe multi-agent behavior in 15 frontier language models.
What This Paper Is About
Existing AI safety benchmarks mostly test models in isolation, so they miss failure modes that only appear when AI agents interact strategically with one another — coordination failures, conflict, and races to the bottom. The authors argue that a small, principled set of game-theoretic structures can capture the dominant strategic tensions behind these multi-agent risks, and they operationalize this by mapping over 1,600 entries from the MIT AI Risk Repository onto canonical games and generating high-stakes scenarios from them. The goal is to measure how often frontier models choose collectively harmful actions in those scenarios, understand why, and test whether institutional "mechanism design" interventions can steer them toward better outcomes.
Key Contributions
- GT-HarmBench benchmark. The first benchmark, per the authors, to evaluate multi-agent LLM safety across canonical strategic structures grounded in real-world high-stakes scenarios: 1,535 scenarios spanning Prisoner's Dilemma, Chicken, Battle of the Sexes, Stag Hunt, Coordination, and No Conflict.
- Demonstration of a reliability gap. Across 15 frontier models and 1,535 scenarios, models fail to achieve the socially optimal choice in 38% of high-stakes cases.
- Characterization of framing, order and reasoning effects. The paper quantifies how explicit payoff cues, option ordering, and different styles of chain-of-thought reasoning shift model behavior toward or away from socially optimal play.
- Mechanism design interventions. Five classical mechanisms (Message, Contracts, Mediator, Penalties, Payments) are tested in 4 prompt styles each (20 variants) across 8 models, improving utilitarian accuracy by 14–18%, with mediation performing best.
Main Findings
- Overall failure rate: Averaged across 15 frontier models and 1,535 scenarios, models achieve socially optimal (utilitarian-maximizing) outcomes in only 62% of cases; the failure rate in high-stakes scenarios is 38%.
- Worst in Prisoner's Dilemma: Mutual cooperation occurs in only 44% of Prisoner's Dilemma scenarios, the lowest welfare of any game type studied. Chicken is more prosocial, with both agents cooperating in 80% of cases.
- Coordination failures even with aligned incentives: In Battle of the Sexes, models converge on the same option only 48% of the time without communication; in Stag Hunt, models vary widely in selecting the risky cooperative action.
- Model ordering and capability mismatch: Aggregated performance ranks Anthropic models highest on average, followed by Meta, then OpenAI, then Google, Qwen, DeepSeek and Grok. The paper reports no clear monotonic relationship between standard capability proxies and achieved social welfare. Table 2 shows wide per-game spread — for example Grok 4.1 Fast scores 0.03 on Prisoner's Dilemma and Claude 4.5 Opus scores 0.98, while both reach 1.00 on No Conflict.
- Explicit payoffs shift models toward selfishness: Adding explicit numerical payoffs to the naturalistic scenario raises Nash equilibrium accuracy by +6.20% on average but lowers utilitarian accuracy by -4.06%, suggesting that surfacing the game-theoretic structure activates more self-interested reasoning.
- Order effects undermine coordination: On the Coordination game, models reach 87% baseline accuracy versus 50% for random choice, indicating use of natural focal points; randomly permuting option order costs an average of 15%, though GPT-5 drops only 5–6%.
- Reasoning patterns predict outcomes: From 12,280 decision traces generated by Claude Sonnet 4.5, Claude Opus 4.5, Qwen 3 30B and DeepSeek v3.2, social welfare reasoning (Utilitarian Δ = 0.07, Rawlsian Δ = 0.11) and safety-oriented reasoning (AI Safety Δ = 0.10) are more prevalent in optimal outcomes, while payoff maximization is strongly associated with suboptimal outcomes (Δ = -0.17).
- Mechanisms help, with a trade-off: Against a baseline of Nash accuracy 0.57 and utilitarian accuracy 0.59, all five mechanisms improve utilitarian accuracy, with gains from +0.13 (Contracts) to +0.18 (Mediator). Messages (+0.03) and Contracts (+0.04) maintain or improve Nash accuracy, whereas Payments, Penalties and Mediator each reduce Nash accuracy by -0.06 — which the authors frame as desirable when Nash equilibria are socially suboptimal.
- Highly uneven mechanism responsiveness: Welfare improvements range from +0.01 (Llama 3.2 3B) to +0.30 (Grok 4.1) and +0.28 (Gemini 3 Pro). Claude Sonnet 4.5 (0.78), Gemini 3 Flash (0.80) and Gemini 3 Pro (0.80) reach the highest absolute utilitarian accuracy across mechanism variants. Grok 4.1 shows +0.30 utilitarian gain alongside -0.10 Nash accuracy; Gemini 3 Pro shows +0.28 with -0.09 Nash degradation.
Methodology in Plain English
The authors start from the MIT AI Risk Repository, which contained 1,612 valid entries at the time of the study. Using GPT-5.1, they classified each entry by which of six canonical symmetric 2×2 games could plausibly capture its strategic tension. The mapping was deliberately inclusive: 604 entries (37.5%) were judged to involve genuine multi-actor strategic interaction, producing 1,816 (risk, game) pairs, a mean of 3.01 games per strategic risk.
For each pair, GPT-5.1 (high reasoning effort) generated a first-person, high-stakes scenario instantiating that game, including each player's context, action labels, explicit numerical payoffs in [-10, 10], and a risk severity score from 1 to 10. A separate GPT-5.1 pass (medium reasoning effort) scored each scenario 0–10 on contextualization quality and correctness of game structure, retaining only those scoring ≥ 8 on both. The overall pass rate was 84.5% (1,535 of 1,816).
Quality checks included a human classification study on 30 stratified scenarios (5 per game) with no payoff matrix shown: inter-annotator agreement was κ = 0.84 with raw agreement of 86.7% (26 of 30). A mechanical check confirmed that 1,530 of 1,535 scenarios (99.7%) satisfy the canonical ordinal conditions of their target game, and the dataset covers the MIT taxonomy with a total variation distance of 6.43%.
Evaluation uses welfare functions and accuracy: utilitarian welfare (sum of utilities), Rawlsian welfare (minimum payoff) and Nash social welfare (product of utilities), with utilitarian accuracy reported in the main paper because the three largely agree. Both players are played by the same model (self-play); cross-play results appear in an appendix figure. The same scenarios were then re-run with five mechanism-design prompt interventions in four prompt styles (20 variants) across 8 models.
Why This Matters
Impact on research. The paper argues that multi-agent evaluation provides complementary insight to single-agent safety benchmarks, and it offers a standardized testbed (benchmark and code released at https://github.com/causalNLP/gt-harmbench) connecting AI risk taxonomies to formal strategic analysis. Its comparison table positions it against prior work by combining real-world safety grounding (the MIT AI Risk Repository) with mechanism interventions (5), versus narrower prior benchmarks such as SanctSim (1 instance, 1 mechanism), CoopEval (4 instances, 4 mechanisms), and purely abstract game evaluations.
Real-world applications:
- Military escalation, including decisions about lethal autonomous weapons systems and arms-race dynamics.
- Election manipulation, where strategic actors may defect on shared norms.
- Medical malpractice and other high-stakes professional advisory settings.
- Financial markets and cybersecurity, where interacting automated agents can produce collectively harmful outcomes.
Industry relevance. Because the paper evaluates models in a third-party advisory role — recommending actions rather than acting autonomously — its findings are directly relevant to deployment pipelines that use LLMs for decision support in competitive or multi-party environments. The framing and ordering results suggest that prompting choices (whether to show explicit payoff numbers, how to order options) can measurably change whether a deployed system recommends cooperation or defection.
Future Directions
- Beyond symmetric 2×2 games. The authors name asymmetric settings (such as human-AI oversight), sequential or extensive-form games (such as inspection games), multiple-party interactions and coalition formation, and incomplete-information games as natural extensions. They view symmetric games as a foundation whose behavior must be understood before asymmetric results can be interpreted.
- First-person agentic evaluation. Moving from third-party advisory framing to settings where AI systems act as principals or autonomously on behalf of users is described as an important extension left to future work.
- Training-based alignment. The mechanism interventions are context modifications rather than training. Whether better-aligned multi-agent behavior can be elicited through reinforcement learning or supervised fine-tuning on game-theoretic objectives remains open.
- Addressing the Nash–utilitarian trade-off. Mechanisms such as Mediator, Payments and Penalties improved social welfare but reduced Nash accuracy, raising the question of how to design interventions that improve collective outcomes without degrading strategic competence.
Target Audience
AI safety and alignment researchers working on multi-agent systems; game theorists and mechanism-design researchers interested in LLM behavior; evaluation and red-teaming teams at frontier labs; and policy or governance analysts who need an operationalized, scenario-grounded account of multi-agent AI risk. Readers with only a passing knowledge of game theory can follow the main results, while the appendices (duality, equilibrium characterizations, prompts and rubrics) serve those who want to reproduce or extend the benchmark.
Authors’ abstract
Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. The benchmark and code are available at https://github.com/causalNLP/gt-harmbench.