Research
The Emergence of Complex Behavior in Large-Scale Ecological Environments
Summary: The Emergence of Complex Behavior in Large-Scale Ecological Environments Overview Research area: Multi-Agent Systems / open-ended evolutionary simulation / neuroevolution and artificial ecolo

- arXiv
- 2510.18221
- Published
- 2025-10-21
- Authors
- Joseph Bejjani, Chase Van Amburg, Chengrui Wang, Chloe Huangyuan Su, Sarah M. Pratt, Yasin Mazloumi, Naeem Khoshnevis, Sham M. Kakade, Kianté Brantley, Aaron Walsman
AI summary
Summary: The Emergence of Complex Behavior in Large-Scale Ecological EnvironmentsOverview
Research area: Multi-Agent Systems / open-ended evolutionary simulation / neuroevolution and artificial ecology (arXiv:2510.18221v3 [cs.MA], published 2025-10-21, revised 12 Dec 2025).
Technical level: Intermediate to Advanced. The paper assumes familiarity with multi-agent reinforcement learning formalisms (MDPs, stochastic games), neuroevolution, and GPU-accelerated simulation with JAX, though its central argument is conceptual rather than mathematical.
Scope: The paper uses a new JAX-based multi-agent simulator to study how physical environment size and population size shape the behaviors that emerge when tens of thousands of objective-free neural network agents evolve through reproduction, mutation, and selection.
What This Paper Is About
Most embodied AI research optimizes a small number of agents against a fixed reward signal in a relatively small world. This paper asks the opposite question: if you place a very large population of agents with no rewards and no learning objective into a physically large environment, and let them survive and reproduce purely through mutation and selection, what complex behaviors appear on their own? The authors build a simulator large enough to test this, scaling to populations of over 60,000 agents in worlds of over 1,000,000 grid cells, and then measure which behaviors emerge, and at what scale.
Key Contributions
- A new open-ended ecological simulator. The authors introduce a JAX-based multi-agent environment, described formally as an "Ecological Game" (EG) and, since agents cannot see the whole state, a partially observable Ecological Game (POEG). Unlike a stochastic game, it removes the reward function and adds a population function that tracks which agents are alive and who their parents are, so births and deaths can be inferred.
- Scaling to populations and environments far larger than prior work. The simulator reaches populations of over 60,000 agents, each with its own evolved neural network policy, in worlds of over 1,000,000 grid cells, run on Nvidia H100 and A40 GPUs.
- Evidence that scale and sensing change what evolves. Through controlled comparisons of sensor configurations (resource sensing only, plus compass, plus vision) across five grid sizes and six terrain types, the paper shows that certain behaviors appear only reliably at sufficiently large scales and that larger scales increase the stability and consistency of emergent behavior.
- A public code release. Experimental code is available at https://github.com/jbejjani2022/ecological-emergent-behavior.
Main Findings
- Long-range "mining" requires a compass and reliably appears only at scale. In the Beach world, agents with a compass (RC) evolved to travel onto land, harvest the biomass deposited there by dead ancestors, and return to water. With four seeds per configuration, mining appeared in 4/4 seeds at 1024×1024, 4/4 at 512×512, 4/4 at 256×256, 3/4 at 128×128, and 2/4 at 64×64. Without a compass (R), mining appeared in only 1/4 seeds at 1024×1024 and 0/4 at every smaller size.
- Population stability tracks scale and compass access. (RC) populations went extinct before 2M steps in 0/4 seeds at 1024×1024, 512×512, and 256×256, 1/4 at 128×128, and 3/4 at 64×64. (R) populations went extinct in 4/4 seeds at 512×512, 256×256, 128×128, and 64×64 (0/4 at 1024×1024).
- A concrete ecological failure mode without a compass. In a representative 256×256 Beach run with (R) agents, water-dwelling agents that wandered onto land could not find their way back and died, depositing biomass inland. Free biomass on dry land nearly doubled by 50,000 steps, the population shrank as water resources were depleted, almost all biomass had moved to land by 125,000 steps, and the final agent died around 185,000 steps. The (RC) version of the same experiment stabilized for the full 2,000,000 steps after the miners cleared the beach.
- Terrain matters. At 512×512, (RC) agents adapted more easily in the Island (4/4 mining, 0/4 extinction) and Isthmus (4/4 mining, 0/4 extinction) terrains than in Lake (2/4 mining, 2/4 extinction) and Channel (3/4 mining, 1/4 extinction). The authors hypothesize this is because walking straight east or west always leads back to water on the island and isthmus, but no such simple rule exists for the lake and channel.
- Vision improves foraging efficiency. In the Ocean world, vision agents (RCV+A) at 1024×1024 averaged a population of 63,198 with biomass utilization .371, move actions .293, and eat actions .504, versus blind agents (RC+A) at 22,788, .146, .425, and .386. The pattern holds at every tested scale (512×512, 256×256, 128×128): vision agents move less and eat more.
- Vision agents attack rarely but precisely. At 1024×1024, vision agents averaged .015 attacks per agent with .732 homicides per attack; blind agents averaged .065 attacks with .239 homicides per attack. The paper reports that agents with vision maintain a successful kill rate of approximately 75%.
- Some behaviors only appear at sufficient scale, and larger scale means less variance. The authors state that some behaviors appear only in sufficiently large environments and populations, and that larger scales increase the stability and consistency of these emergent behaviors. They note that vision runs had "much more stable dynamics" with less variation over time, which they describe as unexpected.
- Survival analysis confirms the scale and compass effects. Kaplan–Meier estimators were fit on 64 runs each for 128×128 (with and without compass) and 32 runs each for 512×512 (with and without compass), with 95% confidence intervals. A Cox proportional hazards model found extinction rate at 128×128 roughly 25 times greater than at 512×512 (HR = 0.04), with compass access producing a similar effect (HR = 0.03); for mining emergence, larger worlds accelerated it (HR = 6.07) and compass use accelerated it dramatically (HR = 34.59). Concordance was 0.80 for extinction and 0.85 for mining.
- A reported contradiction worth flagging. The text states that "mining is more frequent and happens sooner at smaller grid sizes," which runs counter to the seed-level results in Table 1 and to the Cox model result that larger worlds generate mining faster. The provided content does not reconcile these statements.
Methodology in Plain English
The authors built a large grid-world simulation where digital agents must collect three resources — water, energy, and biomass — to stay alive and reproduce. Water flows downhill across a height map; biomass stays where it is but energy slowly grows on free-standing biomass. Agents have health points that decay if they overspend resources, recover slowly if they spend energy and water, and decline with age, so agents eventually die of old age if nothing else kills them. Agents can attack, killing everything in the 3×3 square in front of them and then consuming the dropped resources, which functions as a form of predation.
Each agent's "brain" is a small, memoryless multilayer perceptron, roughly 10k–25k parameters depending on which sensors are active. Sensor readings are flattened, normalized to [-1, 1], passed through separate linear layers with output dimension 64, summed, and fed into a 2-layer MLP with hidden dimension 64 and ReLU nonlinearities. The action space is discrete, sampled from a softmax whose temperature also evolves. Weights use bfloat16 half precision to allow larger populations and faster runs.
Reproduction is single-parent: an agent that has accumulated enough excess biomass creates a child directly behind it, whose policy is a copy of the parent's weights with Gaussian noise of standard deviation 3×10⁻² applied, plus a small inherited resource endowment. Initial populations use Kaiming initialization with zero bias vectors.
Agents can be given different sensor suites: internal and external resource sensing only (R), plus a compass (RC) that reports global direction as a 4-way one-hot vector, plus vision (RCV) that provides a 7×7 top-down rendered patch with three color channels and a fourth channel showing local elevation difference. "+A" denotes that attacking is enabled. Agents also carry a 3-channel color trait that mutates like the weights, which in principle lets vision-equipped agents distinguish types of other agents.
Experiments compare these configurations across six terrains (Ocean, Beach, Island, Lake, Isthmus, Channel) and five grid sizes (64×64, 128×128, 256×256, 512×512, 1024×1024), with initial populations of 128, 512, 2048, 8192, and 32768 respectively — initial population scales proportionally with area. Runs last 2M environment steps or until extinction. Wall-clock time per seed ranges from under an hour for 64×64 to approximately 10 hours for 1024×1024 on a single H100 GPU. Mining is counted as a drop of at least 10% in free biomass on dry land from its historical high point.
Why This Matters
Impact on research. The paper argues that scale does not merely improve a single model but transforms what an entire ecosystem produces, drawing an explicit analogy to the way reasoning abilities emerge in individual language models only beyond a certain size threshold. It positions ecology itself as an instrument for producing machine intelligence, and it provides a concrete demonstration that rich behaviors can be evolved in large populations within a reasonable hardware budget without any explicit objective. It also extends and contrasts with prior multi-agent emergence work — Neural MMO (up to 128 concurrent agents, no generational inheritance), Hamon et al. (2023), and Lu et al. (2024, at most 256 agents) — by combining natural population growth and shrinkage, objective-free evolution, and generational inheritance at much larger scale.
Real-world applications:
- Artificial life and evolutionary computation: a testbed for studying open-ended evolution at scales where collective ecological patterns can appear, without the cost or risk of field experiments on wild populations.
- Multi-agent systems research: an objective-free environment for observing how competition, resource scarcity, and predation shape population-level dynamics in a large agent population.
- Hardware-accelerated simulation engineering: a demonstration that JAX-based simulation with low-precision weights can push agent counts into the tens of thousands, informing how future large-scale multi-agent simulators are built.
- Ecology and conservation modeling: a sandbox for exploring how resource redistribution by individuals can alter an environment in ways that feed back on the whole population — the biomass-transfer-to-land dynamic in the Beach world is exactly this kind of coupled failure.
Industry relevance. The work speaks to organizations building large-scale multi-agent simulation for games, robotics, and synthetic data generation, and to anyone exploiting GPU acceleration to scale simulated populations rather than single policies. Its emphasis on genetic inheritance as a mechanism for producing diverse behavior, without reward engineering, is relevant to teams exploring alternatives to reinforcement learning for open-ended agent behavior.
Future Directions
- Reconciling the survival-analysis and seed-count results. The text says mining happens sooner at smaller grid sizes while Table 1 and the Cox model indicate the opposite. Clarifying this is a prerequisite for the paper's central scale claim.
- Extending the ablations into the main text. The paper defers ablations on initial population size (F.1), resource rules (F.2), mutation rate (F.4), agent network configuration (F.5), vision range (F.6), attack strength (F.7), and agent health points (F.8) to the appendix. Quantifying how sensitive the emergent behaviors are to each of these is a natural next step, as is the experiment in F.3 verifying that dynamics in larger environments are not simply due to self-averaging.
- Richer evolutionary mechanisms. The authors deliberately use only mutation and single-parent reproduction, noting that crossover and more complex approaches are possible. Whether crossover or memory in agent policies produces qualitatively different emergent behaviors at these scales is untested here.
- Explaining the mechanisms behind the observed effects. The paper notes confounding factors in interpreting why blind agents sometimes show high kill rates — clusters of agents forming in one location, which almost never happens with vision-equipped agents. Whether vision-equipped agents are actively learning to avoid each other, or whether the effect is a side effect of population structure, remains an open question.
Target Audience
Researchers in multi-agent systems, artificial life, and neuroevolution; engineers building large-scale GPU-accelerated agent simulations; and machine learning scientists interested in questions of scale and emergent capability outside the standard supervised and reinforcement learning paradigms. Readers looking for a novel optimization algorithm or a benchmark-beating policy will not find one here — the value is in the ecological framing and the scaling demonstration. The paper is most useful to readers comfortable with evolutionary computation and multi-agent formalism, and the appendix (only partially included in the provided content) carries much of the quantitative sensitivity analysis.
Authors’ abstract
We explore how physical scale and population size shape the emergence of complex behaviors in open-ended ecological environments. In our setting, agents are unsupervised and have no explicit rewards or learning objectives but instead evolve over time according to reproduction, mutation, and selection. As they act, agents also shape their environment and the population around them in an ongoing dynamic ecology. Our goal is not to optimize a single high-performance policy, but instead to examine how behaviors emerge and evolve across large populations due to natural competition and environmental pressures. We use modern hardware along with a new multi-agent simulator to scale the environment and population to sizes much larger than previously attempted, reaching populations of over 60,000 agents, each with their own evolved neural network policy. We identify various emergent behaviors such as long-range resource extraction, vision-based foraging, and predation that arise under competitive and survival pressures. We examine how sensing modalities and environmental scale affect the emergence of these behaviors and find that some of them appear only in sufficiently large environments and populations, and that larger scales increase the stability and consistency of these emergent behaviors. While there is a rich history of research in evolutionary settings, our scaling results on modern hardware provide promising new directions to explore ecology as an instrument of machine learning in an era of increasingly abundant computational resources and efficient machine frameworks. Experimental code is available at https://github.com/jbejjani2022/ecological-emergent-behavior.