Skip to content
AI.info

Research

MiniCorp: The Last Mile of the AI Agent Firm

Overview Research area: Artificial intelligence — multi-agent LLM systems, agent-based economic simulation, and enterprise agent training/evaluation. Technical level: Advanced. The paper combines LLM

MiniCorp: The Last Mile of the AI Agent Firm
arXiv
2610.05912
Published
2026-10-05
Authors
Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin, Dylan Zhang, Yuxuan Lu, Qi He, Dakuo Wang, Kai-Wei Chang

AI summary

Overview

Research area: Artificial intelligence — multi-agent LLM systems, agent-based economic simulation, and enterprise agent training/evaluation.

Technical level: Advanced. The paper combines LLM agent architecture, marketplace mechanism design, econometric validation practice, and agent-based modeling conventions (pattern-oriented modeling, operational validation).

Scope: MiniCorp is a checkpointable office simulator in which a persistent firm of LLM agents runs a simulated e-commerce company inside an evolving external market, producing longitudinal and counterfactual enterprise data for training and evaluating agents.

What This Paper Is About

Training agents to run a business requires longitudinal records that connect decisions to their context and downstream consequences, but such records are scarce, expensive, privacy-restricted, and frozen — they contain only the decisions that were actually made, not the outcomes of alternatives. MiniCorp addresses this by simulating a company and the market it operates in, so that enterprise data can be generated at scale and the same situation can be replayed under different decisions.

The goal is twofold: to provide an environment for studying whether agents can collectively run a company, and to serve as a scalable source of longitudinal and counterfactual enterprise data for agent training and evaluation, with the simulator's external world validated against market-response patterns documented in empirical studies.

Key Contributions

  1. MiniCorp, a closed-loop office simulator with an empirically evaluated external world. It couples an evolving external market with a persistent agent firm, producing continuous streams of enterprise data that link organizational communications and decisions to their market consequences. Using e-commerce as a demonstration, the paper presents a general approach to constructing industry-specific external worlds by integrating established economic mechanisms, empirical evidence, and real-world data into explicit market processes with adaptive participants, plus a methodology for end-to-end fidelity evaluation.

  2. A scalable and efficient enterprise data engine. Because the company runs continuously, it leaves behind the record a real company leaves behind (messages, email, meetings, tickets, documents, and similar artifacts) at a volume and speed no privacy-constrained corpus can match, with full provenance linking every artifact to the market state and the decision that produced it.

  3. Checkpointable counterfactual replay. Because its state is checkpointable, the same situation can be replayed under different decisions, yielding the counterfactual pairs a static archive cannot supply.

  4. Fidelity evidence for the integrated external world. The paper reports outcome-level and pathway-level validation tests showing that the integrated simulator reproduces market-response patterns consistent with empirical studies, including response chains and effects that were not direct calibration targets.

Main Findings

  • Price response matches empirical magnitude. A 10% price reduction and a 10% price increase yielded absolute arc elasticities of 2.9 and 3.2, respectively. Demand increased when price fell and decreased when price rose, and the observed elasticities were close to the empirical benchmark of Bijmolt et al. (2005).

  • Temporary discounts produce persistent post-discount sales. When prices were reduced by 20% for four weeks and then restored to their original levels, units increased by 71% during the discount, remained 31% above the control for the first four weeks after restoration, and remained 11% above it for the following six weeks. Nearly half of the incremental units were sold after the discount ended. The pathway traced was: temporary price reduction → higher sales → higher recent sales velocity → more appearances in organic search results → more impressions → persistently higher sales after price restoration.

  • Temporary stockouts produce the mirror-image persistent loss. The sales loss persisted after a two-week stockout: 59% of the cumulative unit shortfall occurred after restocking, and sales remained 36% below the control in week 26. The pathway traced was: temporary stockout → no sales while inventory was unavailable → lower recent sales velocity → fewer appearances in organic search results after restocking → fewer impressions → persistently lower sales.

  • Advertising generates delayed organic gains with cost and competitive trade-offs. Ad-attributed units increased by 23–32%, while the organic gain grew from 0.5% to 5.8%. Each additional advertising dollar returned $0.53 in contribution within the observation window, and 84% of incremental units came from competitor displacement rather than market expansion. The pathway traced was: higher advertising expenditure → more paid sales immediately → higher recent sales velocity → more appearances in organic search results → more organic impressions → a delayed increase in organic sales.

  • Agents coordinate across roles and adapt to market feedback. The abstract reports that experiments show agents coordinating across roles and adapting their decisions to market feedback.

  • Long-term strategic guidance sustains exploration. With explicit long-term strategic guidance, agents sustain advertising exploration despite weak early returns.

  • Organizational realism is measured along three axes. The paper measures the provenance of each work item (whether the queue is driven by colleagues or by an agent's own plan), the process signature of each work category (whether different kinds of work follow different routines), and the delegation structure (which pairs of roles exchange work, and why). The truncated content states only the opening claim that work is generated by the organization rather than by per-agent scripts; the remaining RQ2 results are not reported in the available text.

  • Robustness is not yet established. The paper states that robustness across independently initialized worlds remains to be evaluated.

Methodology in Plain English

The simulator is built around two stateful systems that interact in fixed-length cycles called periods (a week in the e-commerce demonstration).

The external world is an autonomous economy that runs whether or not the agent firm acts: shocks fire, in-flight orders arrive, competitors move, and demand shifts across categories. It models consumers, suppliers, logistics providers, competitors, and a platform, plus macroeconomic conditions. Formally, the next world state is drawn from a conditional distribution given the current state, the firm's decisions, and exogenous events such as tariff changes and freight-rate spikes. Crucially, the world is parameterized by latent structural parameters the agent firm never observes — such as consumer preference weights over price and quality, or a supplier's true defect propensity — so the firm must infer how the market works from observed business data rather than looking up the answer.

Building the marketplace. Rather than mapping an action directly to sales, the simulator models the full response chain: search demand, query–listing matching, organic ranking, search-results pages, click and conversion behavior for each query–SKU pair, sponsored-search auctions, and competitor responses. Search demand is built from 1,000 queries — 316 relatively common queries collected from Amazon autocomplete and 684 more specific long-tail queries constructed from product attributes — with traffic assigned by Zipf's law. ESCI examples are used as few-shot demonstrations for an LLM that classifies each product category as exact, substitute, complementary, or irrelevant to each query. Incumbent sellers are given operating histories through a 26-week warm-up period before the agent firm enters, a state the paper calls simulator-generated prehistory. Sponsored search uses a multi-slot second-price auction for four sponsored positions, with budgets depleting during the week. Competitors are active rather than static: they adjust prices and advertising based on their own observed sales, margins, sales rank, and advertising reports, so the competitive environment evolves endogenously.

The agent firm is a set of standing roles, not agents assembled for a task. Each agent holds one role for the whole run; the structure exists first and work is routed into it. Roles are defined by the business data they own and the set of decision types they may propose — no process, evidence standard, document format, or escalation rule is specified. Agents communicate by email and chat with no central scheduler, each managing its own message queue. All roles see the same sanitized business data and differ in what they may do, not what they may know. Actions must come from a fixed catalogue so the world model can parse and execute them.

Interaction cycle. At the start of a period, the world exports a filtered copy of its business data through several masks: whole tables are dropped when they hold latent parameters, internal causal records, or answers; protected columns are removed; a final schema audit fails the entire export if anything protected survives. The export is raw and unanalyzed, so agents must define and compute their own derived measures. Approved decisions and exogenous events are applied together at the period boundary, and a decision approved in one period takes effect only in the next, preventing the firm from observing and rewriting an outcome within the same period. A week ends when all employees are idle with no message in flight, or when a time/step backstop is reached, which keeps message ordering reproducible and independent of LLM latency.

Evaluating fidelity. Because market conditions and numerical magnitudes vary across products, firms, platforms, and time periods, the authors treat exact numerical agreement as supporting evidence only, and use direction, shape, temporal dynamics, distribution, and trade-offs as primary criteria. They follow pattern-oriented modeling, operational validation, and empirical-validation frameworks for agent-based models, running controlled interventions from the same checkpoint and, where applicable, with the same random draws, then tracing the full response trajectory. Emphasis is placed on response chains and outcomes that were not direct calibration targets.

Experimental setup. The agent substrate is powered by LLM APIs of the frontier model GPT-5.6-Sol.

Why This Matters

Impact on research. The paper reframes progress toward enterprise AGI as a third stage beyond chatbots and task-completing autonomous agents: organizational autonomy, the capacity of a system of agents to set strategic direction, decide under uncertainty, coordinate across roles, and keep an organization alive and profitable over time. Existing benchmarks (τ²-Bench, SWE-bench, GAIA, OSWorld) evaluate bounded tasks with predefined objectives and verifiable outcomes, and even multi-agent benchmarks typically assemble agents around a task and discard them when it ends. MiniCorp targets the missing combination: a persistent organization, an evolving external market, co-evolution between them, and checkpointable replay. It also names and takes seriously the failure mode of simulator overfitting, where optimization reinforces regularities that exist in the simulator but not in the real environment.

Real-world applications:

  • Enterprise agent training and evaluation. Generating longitudinal, provenance-linked enterprise data without the privacy and compliance constraints that make internal corporate records unavailable for training.
  • Counterfactual business analysis. Replaying a recorded business situation under alternative decisions to study what-ifs that static archives cannot answer.
  • Consulting and strategic decision support. Using the validated external world as a sandbox to test pricing, advertising, and inventory decisions before committing them in a real market.
  • Organizational design research. Studying how roles, authority, and delegation structures affect outcomes in a persistent simulated company.

Industry relevance. The paper notes that Google was selected as the winning bidder in a bankruptcy auction with a $10 million offer for Spirit Airlines' internal business data, including employee emails, Microsoft Teams messages, spreadsheets, calendars, and operational records, which it planned to use for product development and AI training. It also cites the Enron corpus as an example of fragmented public records. This framing positions simulated enterprise data as an alternative to expensive, restricted, and incomplete real archives. Because deployment is conditioned on the simulator resembling the market it represents, and because the authors explicitly flag robustness across independently initialized worlds as still unevaluated, the industrial claim depends on continued fidelity validation.

Future Directions

  1. Robustness across independently initialized worlds. The paper states this remains to be evaluated for the default simulator configuration, even though the fidelity tests passed.

  2. Refining response pathways identified as missing or incorrect. The end-to-end fidelity evaluation is designed to identify missing or incorrect response pathways that require further refinement, implying an iterative construction-and-validation loop.

  3. Correcting the author-identified limitations of archives. The paper identifies two gaps in historical archives — fragmented context and frozen single trajectories that cannot answer counterfactual questions — as the problems MiniCorp is built to solve; extending counterfactual generation to broader settings follows naturally.

  4. Extending beyond the e-commerce demonstration. The paper describes its approach as a general method for constructing industry-specific external worlds by integrating economic mechanisms, empirical evidence, and real-world data, with e-commerce serving as the running example.

Explicitly not reported in the available content: the remaining RQ2 organizational-realism results on process signatures and delegation structure, results for any research questions beyond RQ1 and RQ2, the duration of a full run, the number of roles in the agent firm, and the performance of any comparison or baseline models.

Target Audience

Researchers and practitioners working on LLM-based multi-agent systems, agent evaluation and long-horizon benchmarks, agent-based economic simulation, and enterprise AI deployment. It is also relevant to economists and organizational researchers interested in computational models of firms and markets, and to industry teams that need longitudinal enterprise data for training or evaluating agents but cannot access proprietary corpora. Readers without background in multi-agent systems or market mechanism design will need to work through the methodological sections carefully.

Authors’ abstract

The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They cannot show the outcomes of alternative decisions. We introduce MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale. Using an e-commerce company as a demonstration, MiniCorp connects two interacting worlds. The external world models customers, dynamic competitors, and market mechanisms. The internal world consists of agents that observe events, discuss their options, and make strategic decisions. These decisions have lasting effects on the market, and the resulting feedback informs the firm's later decisions. As the firm and market interact, MiniCorp continuously records the agents' communications and decisions. These records preserve the information available at the time and the business results that followed. Checkpointing allows the same situation to be replayed under different decisions, providing comparisons unavailable in static archives. We evaluate end-to-end fidelity against patterns reported in empirical studies of real markets. These evaluations provide agents with realistic market feedback and reduce the risk that they learn to exploit flaws in the simulator. Our experiments show agents coordinating across roles and adapting their decisions to market feedback. With explicit long-term strategic guidance, they also sustain advertising exploration despite weak early returns. MiniCorp thus provides an environment for studying AI-run companies and a scalable source of longitudinal and counterfactual enterprise data for agent training and evaluation.

Read the original paper