Research
SWE-Game: Can Coding Agents Build the Games We Want?
Overview Research area: Evaluation of AI coding agents on long-horizon, executable software tasks, specifically game development in Godot and Unity. Technical level: Intermediate. The task definitions

- arXiv
- 2609.33678
- Published
- 2026-09-27
- Authors
- Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang
AI summary
Overview
Research area: Evaluation of AI coding agents on long-horizon, executable software tasks, specifically game development in Godot and Unity.
Technical level: Intermediate. The task definitions and scoring are explained clearly, but the evaluation machinery (instrumentation bindings, certified input replay, deficiency ceilings in rubrics) is detailed and assumes some familiarity with game engines and agent benchmarks.
Scope in one sentence: The paper introduces SWE-Game, a 247-task benchmark built from 41 executable reference Godot games, and uses it to measure whether six coding agents can build, complete, repair, and port games that match an intended, reference-defined gameplay.
What This Paper Is About
Existing game-development benchmarks tend to evaluate generated games with coarse rule checks or model-judge scores, which can miss cases where a game looks right but behaves wrong (for example, spikes that are visible but harmless). The paper's goal is to connect what a task asks for to what the finished game actually does at runtime, by grounding every task in a playable reference game rather than only in text, images, or video. It then measures both behavioral correctness through executable checks and presentation quality through game-specific visual rubrics.
Key Contributions
- A benchmark grounded in executable reference games. 247 tasks are derived from 41 Godot reference games across 13 gameplay categories in 2D and 3D (24 in 2D, 17 in 3D), covering five development activities.
- Runtime evaluation across independently written implementations. A shared instrumentation interface maps different object names, scene hierarchies, and code to semantic roles, standardized actions, and observable state, so evaluator-owned drivers and probes can run input replays and engine-state checks on submissions that do not match the reference project structure.
- A systematic study of agent capability and evaluator validity. Six models are evaluated across all 247 tasks, reviewed submissions are categorized by implementation problem type, and automated judgments are compared against human labels for both functional and visual assessment.
- Behavior-dependent scoring rules. Certified routes earn no credit if a matched no-input control also reaches the goal, and agent-authored feature demonstrations count only if the demonstrated outcome does not occur in a no-input run of matched duration.
Main Findings
-
One model leads on every task type. Opus5 achieves the highest overall score in all five task types, ranging from 50.38 in Brief-to-Game generation to 83.46 in Bug Repair. Below Opus5, rankings shift by task: Grok4.6 is second in Brief-to-Game, and GPT-5.6 Luna is second in the other four tasks.
-
Construction tasks remain difficult. Best overall scores stay below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Every model scores higher on GDD-to-Game implementation than on Brief-to-Game generation, although the two tasks differ in both inputs and scoring criteria.
-
Content realization is the weakest component in Skeleton Completion. Content scores range from 11.97 to 36.93 across the six models.
-
Repair performance varies mainly in restoration. In Bug Repair, Restoration spans 35.27 to 90.18 across models, a wider range than Retained routes (96.00 to 99.51) or Preservation contracts (75.51 to 93.17). Both restoration and preservation contribute to the repair score.
-
Porting shows a structure-versus-presentation gap. GPT-5.6 Luna scores 82.87 in Structure, 51.44 in Playability, and 31.60 in Visual, while Opus5 obtains 90.53, 65.88, and 58.40 in those same components.
-
Requirement omissions and gameplay logic errors dominate failures. Requirement omissions are the most common categorized problem for GPT and Opus, at 56.4% and 36.9% respectively. Engine-integration errors account for another 17.3% of GPT's categorized problems. For GLM, gameplay logic errors are the largest category at 52.2%.
-
Executable checks agree with humans more closely than a video-based judge. On 1,200 human-labeled assertions from 100 agent-built games, executable checks reach 92.59% balanced accuracy versus approximately 78.41% for the video-based VLM judge. Executable checks correctly accept 411/452 correct assertions and detect 705/748 defective ones, with an 18.85-percentage-point advantage in defect detection (94.25% versus 75.40%); correct acceptance is 90.93% versus 81.41%.
-
Visual rubrics correlate strongly with human ratings. Rubric-based VLM scores on 200 gameplay clips reach a Spearman correlation of 0.829 with human ratings, with a mean absolute error of 0.103 on a [0,1] scale and a within-clip standard deviation of 0.034 across repeated judgments.
-
Reference video mainly improves presentation, not mechanics. Removing the reference video lowers the VLM total by 5.60 points for GPT-5.6 Luna and 2.23 for Opus5, and lowers Visual Presentation by 6.48 and 9.00. Other component changes are mixed: with video, GPT-5.6 Luna gains 10.30 in Gameplay Realization and 10.14 in Guidance and Feedback, while Opus5 changes by -3.94 and +1.33; Content Composition is lower with video for both models, by 6.01 and 9.93.
-
Resource use differs sharply between comparable scores. In Brief-to-Game, Opus5 attains the highest score with 8.49M input tokens, 127.34k output tokens, and 111.37 tool calls, compared with 2.56M, 34.05k, and 57.15 for GPT-5.6 Luna.
Methodology in Plain English
The team first built 41 reference games in Godot (24 in 2D, 17 in 3D), starting from licensed public asset packs and writing a game design document for each that fixes the objective, controls, core mechanics, success and failure conditions, and content scope. Levels follow an introduce, vary, combine progression, and designs were revised through runtime checks and human playtesting. Across all references, the games contain 168,335 lines of gameplay GDScript, 1,587 scenes, 22,140 scene-tree nodes, and 138 gameplay recordings; gameplay code ranges from 1,153 to 11,024 lines per project (median 3,367) and scenes range from 7 to 108 (median 32).
From these references, five task modes were derived. Three are construction modes differing in what is supplied: Brief-to-Game (a short request, requiring the agent to author its own GDD), GDD-to-Game (the full reviewed design), and Skeleton Completion (an unfinished project with missing functionality to integrate). Bug Repair supplies a project that launches but misbehaves, plus a player-facing symptom report that does not name files or lines, with 83 cases and faults including state not reset across levels and partially migrated mechanisms. Godot-to-Unity Porting supplies the Godot source, GDD, assets, reference video, and a target Unity interface, and requires a buildable Unity project preserving controls, mechanics, and progression. Each mode has one task per game except Bug Repair, giving 247 tasks total.
To evaluate submissions that legitimately use different object names and scene hierarchies, games declare semantic roles in files such as gb_levels.json, tag objects with groups such as gb_player and gb_hazard, and bind task-relevant properties like health, progress, and score. The evaluator, not the submission, owns the checks and the expected outcomes. Evaluation combines engine-state checks and certified routes (input sequences authored and validated on the reference game, retaining reference action timing and checking predefined goals and intermediate milestones), agent-authored feature demonstrations, and matched no-input controls that invalidate any outcome that happens without player action. Repair cases run identical scenarios on the reference, faulty, and repaired versions, and qualify only if the reference passes, the faulty version fails the designated behaviors, and unaffected checks remain passing. Porting adds project and interface readiness, Unity build, and runtime behavior checks in the target engine. For the construction modes, 41 game-specific rubrics containing 658 items across Gameplay Realization, Content Composition, Guidance and Feedback, and Visual Presentation score recorded frames. Construction scores combine a normalized objective score (ceiling of 85 points) and a visual score using S = 0.85 O + 0.15 V, with visual group weights of 0.10, 0.18, 0.27, and 0.45 renormalized over applicable groups; all results are reported on a 0 to 100 scale.
Why This Matters
For research, the paper shows that benchmark tasks in interactive software should specify behavioral targets that evaluators can execute against, not just artifacts that can be inspected or judged. Grounding evaluation in reference games allows alternative implementations to be accepted while still testing whether required interactions actually work, and it quantifies how much more reliable executable checks are than a video-based judge at detecting defects.
Real-world applications:
- Automated QA for game studios. Runtime checks of mechanics, progression, and input dependence could be embedded in build pipelines to catch regressions such as hazards that no longer damage the player.
- Cross-engine migration support. The Godot-to-Unity porting task models a common production activity: moving an existing project to another engine while preserving controls, mechanics, and progression.
- Agent workflow diagnostics. Categorizing failures into requirement omissions, gameplay logic errors, and engine-integration errors gives concrete targets for improving agent planning, verification, and integration tooling.
- Evaluation design for interactive software. The instrumentation interface pattern, where submissions declare semantic roles and the evaluator owns the tests, is applicable beyond games to any interactive system with implementation freedom.
Industry relevance centers on the gap the paper reports: producing a project that launches and looks plausible is easier for current agents than faithfully implementing the intended mechanics and content, since best construction scores remain below 60 out of 100. The paper also notes that generated projects run in isolated evaluation environments without access to external services, credentials, or private user data, and that copyright, licensing, and attribution requirements should be respected when releasing generated content.
Future Directions
- Expand the reference game collection beyond the current 41 games, which the authors state limits coverage of large projects and long-term development.
- Compare agent frameworks under controlled budgets, since each of the six models was evaluated with a single agent framework and the reported results reflect the combined model-plus-framework capability.
- Extend engine-migration study, as Godot-to-Unity porting is described as an initial study of engine migration.
- Evaluate longer development workflows and broaden human assessment to further refine the visual rubrics and clarify how automated scores relate to player experience.
Target Audience
Researchers and engineers working on coding agents, agent benchmarks, and automated software evaluation will get the most from this paper, since its central argument concerns how to connect task specifications to executable runtime evidence. Game developers and technical QA leads interested in automated playtesting and engine migration will also find the task definitions and failure taxonomy useful. Readers looking for per-model framework details, dollar costs, or comparisons under fixed compute budgets will not find them here, because the paper reports token and tool-call statistics for the recorded runs and notes that accounting rules and tools differ across configurations.
Authors’ abstract
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.