Skip to content
AI.info

Research

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

Overview Research area: Autonomous LLM/VLM agents for software and game engineering; recursive self-improvement; experience internalization through supervised fine-tuning. Technical level: Advanced. T

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
arXiv
2609.39045
Published
2026-09-30
Authors
Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou, Siqi Liu, Aayush Salvi, Yiheng Lin, Ce Zhang, Xiaohan Lan, Jiahui Zhu, Yujie Zhong, Qi She, Biwei Huang

AI summary

Overview

Research area: Autonomous LLM/VLM agents for software and game engineering; recursive self-improvement; experience internalization through supervised fine-tuning.

Technical level: Advanced. The paper assumes familiarity with agentic loops, tool-calling LLMs, benchmark evaluation protocols, and supervised fine-tuning.

Scope: The paper introduces RSIGame, an autonomous agent framework that develops playable games through a local explore–diagnose–improve loop plus a global quality-monitoring loop, and then internalizes successful development experience into the generator model.

What This Paper Is About

Automatically generating a playable game from a natural-language specification is now feasible, but reliably improving that game beyond a merely playable version is not. Naive iterative refinement tends to overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. RSIGame addresses this by organizing autonomous game development into complementary local and global loops, and by transferring verified development experience back into the generator.

Key Contributions

  1. A local explore–diagnose–improve loop. A controller decides what to explore next, an explorer interacts with the executable game to gather behavioral evidence, an editor converts that evidence into concrete edits, and a verifier replays the updated build to check whether the targeted issue was solved without regressions. An evolving checklist accumulates issues, priorities, and verification outcomes across rounds.

  2. A global progress-monitoring and control loop. A game quality monitor tracks accumulated quality across stages and preserves the best checkpoint via a SelectBest operation, while saturation or regression detection decides whether to terminate or open a new stage under sparse high-level guidance. This monitoring is stated to be strictly isolated from the benchmark evaluator, including its scores, rubrics, and feedback.

  3. Training-based experience internalization. Successful generation, planning, and verified improvement trajectories are curated and used for supervised fine-tuning of Qwen3.8-27B, extending self-improvement from context optimization to parameter optimization. The corpus comprises 2,213 generation traces, 2,108 planning traces, and 2,013 verified improvement rounds from 4,003 candidates.

  4. Broad empirical validation. Evaluation spans 140 GameCraft-Bench tasks across 15 game families, two engines (Godot and Phaser), and five generator settings, all under matched development budgets.

Main Findings

  • Consistent gains across generators on Godot. RSIGame improves the frozen initial games by 10.7 to 20.3 Overall points across all five settings, with gains spanning Mechanics, Depth, Visuals, and Art.

  • RSIGame beats prior iterative development under matched budgets. On Godot it outperforms Play2Code by 7.2 to 13.8 Overall points. The paper notes Play2Code brings almost no improvement to the strong Codex+GPT-5.5 initialization and the Qwen3.8-27B variants, whereas RSIGame substantially improves the same frozen games.

  • Small open models overtake much stronger one-shot generators. Qwen3.8-27B with RSIGame reaches 61.38 on Godot, above the Codex+GPT-5.5 one-shot score of 50.26. With iterative development alone (no fine-tuning) it reaches 47.77 on Godot versus Kimi-K2.6's 44.77.

  • Large token efficiency gains. Fine-tuning raises Qwen3.8-27B's one-shot score from 37.07 to 48.22, within 2.0 points of Codex+GPT-5.5, while reducing generation tokens by 11 times (6.41M to 0.57M). With RSIGame, the score rises to 61.38 and total token usage falls 2.6 times (9.13M to 3.49M).

  • Cross-engine transfer. On Phaser, RSIGame improves all three generators by 8.8 to 13.4 Overall points over their frozen bases and achieves the strongest final quality in every setting. Qwen3.8-27B with RSIGame rises to 50.24, past GPT-5.5's 49.44 one-shot score; the fine-tuned Qwen3.8-27B reaches 58.53 on Phaser.

  • Play2Code is stronger on Phaser than on Godot. The category breakdown attributes this to Phaser initializations having weaker Mechanics and Depth but stronger Visuals and Art, leaving more readily improvable functional headroom. RSIGame still produces the best final games.

  • Development-time scaling behaves differently. On a 40-task subset, Play2Code quickly plateaus and often regresses as more rounds are added, while RSIGame converts additional rounds into higher quality on both engines and for both strong (GPT-5.5) and weak (Qwen3.8-27B) initializations. The Global Quality Monitor retains the best state reached so far and closely tracks the oracle best within each budget; saturation-aware stopping reaches comparable quality with fewer rounds.

  • Adaptive development beats fixed scheduling. After a build failure the loop shifts toward improvement, and once visual quality becomes the bottleneck it allocates more rounds to art. Blind pairwise evaluation prefers adaptive development in 59% of comparisons versus 37% for round-robin, with consistent advantages across all four quality dimensions.

  • Agentic verification improves reliability. Evidence-grounded pre-improvement verification raises grounded precision from 58.6% to 72.3% and reduces unsupported targets from 1.93 to 0.50 per round. Post-edit replay-based verification detects 76.2% of unsuccessful improvements and reaches 84.4% balanced accuracy, whereas build-only checking detects 0.0% of these behavioral failures (with 100.0% specificity and 50.0% balanced accuracy).

  • Planning and improvement experience transfer better than generation alone. Training on generation traces alone yields uneven gains, while adding planning and verified improvement experience produces consistent gains across dimensions, most notably Mechanics +14.9 and Depth +14.1, for a +11.1-point Overall improvement before any test-time development.

Methodology in Plain English

The team treats game development the way a human studio would: build, playtest, diagnose, fix, and verify. Rather than letting one model freely rewrite the code, they split the work across specialized agent roles.

Inside each round, a controller looks at the game specification, the current project, an evolving checklist, and any stage-level guidance, and proposes a direction to investigate. An explorer then actually plays the executable game along that direction and records what happens. An editor turns the observed problems into concrete code/asset edits. A verifier replays the updated build to confirm the target problem was fixed and nothing else broke, writing the outcome back into the checklist. The checklist is the shared memory that keeps every agent working from the same picture of what remains broken.

Above this local loop, a global quality monitor periodically compares the current build against a retained best checkpoint and keeps the winner, so later bad edits cannot erode earlier progress. If several consecutive checkpoints fail to improve on the champion, the system treats this as saturation: it either stops and returns the champion, or starts a new stage from that champion under fresh high-level guidance. The paper's runs use a checkpoint interval E = 3 and patience K = 3.

Separately, the team collected successful development trajectories produced with GPT-5.5 (generation), planning traces from successful generations, and improvement traces from running RSIGame with GLM-5.3-Flash. They retained only executable generations and improvements verified through post-edit interaction without breaking previously functional behavior, then supervised fine-tuned Qwen3.8-27B on the curated planning, generation, and improvement trajectories so the model internalizes intermediate decisions, not just final artifacts.

Evaluation follows GameCraft-Bench: each game is packaged with replayable demonstration traces that are replayed and scored against a hidden task-specific rubric over Mechanics, Depth, Visuals, and Art, with Q = BUILD × (0.15M + 0.35D + 0.15V + 0.35A) and BUILD ∈ {0,1} zeroing out non-playable games. Qwen3.8-27B serves as judge, and three independent replay-and-score runs are averaged. All methods start from an identical clone of P₀ with the same per-round tool budget, up to 26 per-round tool calls and 30 development rounds.

Why This Matters

Impact on research. The paper reframes iterative agentic development as a recursive self-improvement problem and argues the bottleneck is not iteration itself but how iteration is organized. It contributes a concrete architecture (adaptive local improvement plus global best-checkpoint tracking and saturation detection), a verification protocol with measured grounding and failure-detection rates, and evidence that development experience can be internalized into model parameters — a step beyond prompt- or context-level refinement.

Real-world applications:

  • Rapid game prototyping. Turning a natural-language spec into a refined, playable build could compress early concept iteration for studios.
  • Automated QA and playtesting. The explorer/verifier structure is a general recipe for finding behavioral bugs in interactive software that compilation checks cannot reveal — the paper shows build-only checking caught 0.0% of the behavioral failures replay-based verification caught.
  • Tooling for small teams and solo developers. Fine-tuning let a smaller open model (Qwen3.8-27B) reach 61.38 on Godot versus GPT-5.5's one-shot 50.26, while cutting generation tokens 11 times — relevant where inference cost and access to frontier models are constraints.
  • General long-horizon agent pipelines. Best-checkpoint retention and saturation-aware stopping are transferable to any agent loop prone to drifting or regressing over many rounds.

Industry relevance. The cost and token accounting (mean billable tokens and cost per task, with RSIGame costing $0.88 to $1.53 per task across reported rows versus Play2Code's $0.99 to $1.75) frames autonomous development in production terms rather than only quality terms. The cross-engine result, where findings on Godot transfer to Phaser, speaks to robustness of the approach across different game engines.

Future Directions

  • Better understanding of guided multi-stage development. The paper states that its study of the optional directed extension, where sparse high-level guidance from a human or stronger model is introduced after autonomous improvement saturates, is limited to a small set of multi-stage cases. When guidance should be invoked, how much is beneficial, and how human and model guidance differ remain open questions.

  • Evaluating dimensions beyond the four reported. Mechanics, Depth, Visuals, and Art do not capture originality, narrative quality, long-term player engagement, or subjective enjoyment.

  • Scaling to larger projects. The experiments focus on relatively compact games developable within practical agent budgets. Extending autonomous recursive development to substantially larger projects with longer horizons, richer assets, and more complex cross-system dependencies is named as an important direction.

  • Continual recursive self-improvement. The paper frames experience internalization as opening the door to continual RSI, where newly acquired development experience is repeatedly internalized to further improve the generator over time. Training-data ablations and the specific effects of repeated internalization rounds are not resolved in the main text.

Target Audience

Researchers and engineers working on LLM/VLM agents, agentic software engineering, automated code repair, and recursive self-improvement will find the loop design and verification measurements most useful. Game AI and game-tooling practitioners evaluating automated prototyping or automated playtesting are the primary applied audience. Benchmark designers and evaluation researchers will care about the cross-engine protocol, the hidden-rubric replay scoring, and the released evaluation artifacts (51,644 scoring files containing per-task scores, judge outputs, and replay reports). It is less suited to readers seeking an introductory treatment of game generation, since it assumes fluency with agent architectures and fine-tuning.

Authors’ abstract

Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.

Read the original paper