Skip to content
AI.info

Research

CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

Overview Research area: Multi-agent systems / agent evaluation benchmarks, specifically long-horizon competitive decision-making by LLM-based agents. Technical level: Intermediate. The economics (logi

CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets
arXiv
2609.34821
Published
2026-09-28
Authors
An Yan, Yu Huo, Zhiwei Shang, Yiran Peng, Chenglin Wu

AI summary

Overview

  • Research area: Multi-agent systems / agent evaluation benchmarks, specifically long-horizon competitive decision-making by LLM-based agents.
  • Technical level: Intermediate. The economics (logit demand, terminal valuation) is explained in the paper, but the core protocol is understandable without deep background.
  • Scope: The paper introduces CEO Arena, a benchmark that evaluates eight LLM-based "CEO" agents competing for 500 simulated days in a shared eight-firm market, and proposes "matched replacement evaluation" to separate an agent's private gains from the effects it has on rivals and on the market as a whole.

What This Paper Is About

Most agent benchmarks measure whether an agent completes its own task well. They rarely measure what an agent does to the other agents around it. CEO Arena asks both questions at once: can an LLM run a company over 500 simulated days of pricing, procurement, marketing, R&D and service decisions, and what happens to the market when it does? The goal is a controlled testbed where private performance and market externalities can be compared directly and attributed to a single agent, rather than only measuring who "wins."

Key Contributions

  1. Matched replacement evaluation. A protocol that compares a focal agent against a reference policy (Rule CEO) in the same seat and the same economic seed, holding the other seven agents' identities and seats fixed while all agents still adapt. This separates private gains, directed externalities on each rival, aggregate externalities, and total market change.
  2. A shared-market testbed. A 500-day, eight-company market with private company information, noisy market signals, resource constraints, delayed feedback, and adaptive rivals. Only active CEO agents decide from the same world version; settlement waits for their decision windows to close. It tests long-term planning under uncertainty, information gathering from noisy signals, adaptation to changing markets, and coordination of business decisions toward a firm's goal.
  3. Empirical evidence on externalities. Across eight LLM-based CEO agents (27 main runs, 26 robustness runs), private gains can accompany contrasting market effects, with repeated-run and reference-policy checks, plus mechanism hypotheses grounded in memory and action traces.
  4. Stability analysis of directed effects. Four of 56 directed agent pairs satisfy |mean| > SD across three main seeds and two seed-11 repeats, and the market-level sign pattern largely survives swapping the baseline policy from Rule to No Action.

Main Findings

  • Most agents lose money. In the main evaluation (nine lineups, seeds 11, 29 and 47, 24 outcomes per agent), only GPT-6 Astra (Mean Score $116.94k), GPT-5.6 Sol ($98.44k) and Qwen 3.8 Max ($8.04k) have positive Mean Scores. Rule CEO averages −$47.30k and beats the remaining five LLM agents: DeepSeek V4 Flash (−$114.50k), Gemini 3.8 Flash (−$119.79k), DeepSeek V4 Pro (−$159.33k), Claude Sonnet 5 (−$314.36k), and Doubao Seed 1.8 (−$434.88k).
  • Bankruptcy is common but not universal. Bankruptcy rates range from 0.0% (Astra, Sol, Qwen, Rule, Gemini) to 4.2% (DeepSeek V4 Flash) and 12.5% (DeepSeek V4 Pro, Claude Sonnet 5), up to 37.5% for Doubao Seed 1.8. Sixteen of 216 firms go bankrupt while 145 end with negative Score, meaning liquidity and profitability are distinct.
  • Private gains need not raise market returns. Astra's private difference averages $116.36k and its market difference $466.68k, both positive in all three seeds. Sol's private difference is positive in all three (mean $131.55k) but its market difference averages −$89.66k and is negative in two of three seeds.
  • Only Astra lifts the market. All three reference markets have negative total Score, so Astra's positive market difference means a smaller aggregate loss than under the matched Rule replacement, not a profitable market. Astra is the only entrant with a positive market difference in 3/3 seeds; DeepSeek V4 Pro, Claude Sonnet 5 and Doubao Seed 1.8 are positive in 0/3.
  • Externalities are directional and not reciprocal. The largest negative directed mean difference is Sol→Sonnet (−$195.66k); the largest positive is Qwen→DeepSeek Pro (+$346.29k).
  • Four of 56 directed pairs are relatively stable. They meet |mean| > SD across three main seeds plus two seed-11 repeats; three retain their sign in all five observations and the fourth in four. The paper states this is a descriptive display criterion, not a significance test.
  • The pattern survives a baseline swap. At seed 11, seven of eight entrants retain the sign of their market Score difference when the baseline changes from active Rule CEO to No Action (which seals an empty decision, preserving standing policies). Private and Others comparisons each retain five of eight signs, and 33/56 directed comparisons retain their mean sign.
  • Removing terminal salvage does not change the sign. Excluding terminal inventory salvage preserves the sign of every entrant's three-seed mean market Score difference and every entrant–rival pair's three-seed mean directed difference.
  • Possible mechanisms. Astra repeatedly trims inventory and unused capacity; Sol sustains greater sales with larger marketing and capacity outlays, then downsizes as demand weakens; Qwen uses explicit reversal rules for prices and budgets. Relative to Rule, Sol adds more own consumer orders than Astra in every seed, and rivals' spending reductions exceed revenue losses with Astra while revenue losses exceed savings with Sol. With Astra, DeepSeek Flash spends less on innovation in every seed; Doubao follows DeepSeek Pro's prices, sometimes adding further discounts.

Methodology in Plain English

The researchers built a simulated economy where eight companies with identical endowments compete in one shared product market for 500 simulated days. The engine controls customers, suppliers and settlement; the CEO agents decide prices, procurement, budgets, projects and bids. The market has three product tiers and three customer segments, and demand is allocated by a logit-style choice rule plus an outside option, so total sales can vary. Each firm can only see its own records and public information, not rivals' private states or future shocks; paid reports give delayed, noisy estimates of market conditions. Decisions are made in windows: a CEO queries records, stages a private draft, and seals it with a tool call — unsealed drafts are discarded and standing policies continue.

To measure an agent's effect rather than just its score, the authors use a leave-one-out design. The pool has eight LLM agents plus Rule CEO, giving nine lineups. Each lineup runs under seeds 11, 29 and 47, producing 27 markets and 216 company outcomes (24 per agent). An agent's terminal Score adds cash plus inventory salvage value and subtracts initial enterprise value; agents are ranked by Mean Score across eight lineups and three seeds. The key comparison is matched replacement: put Rule in agent B's seat while the other seven identities and seats stay fixed, run a fresh world so rivals still adapt, and subtract. This yields a private difference, a directed externality on each rival, a total externality, and a market difference (which equals private plus others, before rounding). Robustness uses two extra evaluations of all nine lineups at seed 11 (18 runs) and eight No Action replacements. All 53 runs reached day 500 and passed model-free replay.

Why This Matters

  • Research impact. It extends agent benchmarking from "did the agent do its task" to "what did the agent do to everyone else," and shows that the two questions can have opposite answers. The matched replacement design gives a concrete recipe for measuring externalities in multi-agent systems, complementing prior work such as Melting Pot's background-population metrics and empirical game-theoretic strategy-profile comparisons.
  • Real-world applications:
    • Automated pricing and promotion systems, where one firm's policy changes rivals' margins.
    • Procurement and inventory policy agents, where stockpiling shifts supply available to others.
    • Marketing and attention-budget agents, where spending captures demand from competitors.
    • Any marketplace where autonomous sellers set prices, in which aggregate consumer or market welfare is not the same as any single seller's return.
  • Industry relevance. Companies deploying LLM agents for business decisions need to know not only whether an agent hits its own KPIs but whether it damages the surrounding market — and the paper's findings suggest the two can diverge sharply. The paper also flags that these results do not establish that the observed policies are suitable for autonomous operation of real organizations.

Future Directions

  • Causal mechanism testing. The authors state that the externality analysis proposes mechanism hypotheses but is not a comprehensive causal analysis; identifying mechanisms will require further targeted interventions and causal-inference analysis.
  • More replication for confidence intervals. The paper describes robustness using means, sample standard deviations and sign consistency, and notes that informative confidence intervals would require more extensive replication.
  • Testing the Rule baseline properly. Rule was selected in a separate 75-world, model-free search with nine firms but evaluated with eight under the same 690 baseline opportunities, which may ease competition. Rule's eight-firm optimality and effect magnitudes remain untested.
  • Less stylized environments. The simulator omits financing, labor markets, regulation, direct inter-agent communication and shared supplier scarcity, and all firms share a product catalog and initial endowments. The authors have not tested whether rankings and externality patterns hold under different demand, liquidity, delay or observation-noise settings, and transfer to real firms remains untested.

Target Audience

Researchers working on LLM-agent benchmarks, multi-agent systems and multi-agent reinforcement learning; economists and mechanism-design researchers interested in simulated market competition; and practitioners who deploy or evaluate autonomous agents for business operations and need to reason about effects that reach beyond the agent's own objective. Readers looking for causal claims or production-ready guidance will find the paper explicitly limits itself to a controlled, stylized testbed with no human participants and with returns that exclude consumer welfare and broader social costs.

Authors’ abstract

Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents' identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals' pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.

Read the original paper