Skip to content
AI.info

Research

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Overview Research area: Software engineering (cs.SE) — evaluation of LLM-based coding agents on cross-repository software change tasks. Technical level: Advanced. The paper is readable without deep sy

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
arXiv
2609.33382
Published
2026-09-27
Authors
Baoyi Wang, Xingliang Wang, Jinyang Wu, Keming Wu, Chen Zhi, Jianwei Yin

AI summary

Overview

Research area: Software engineering (cs.SE) — evaluation of LLM-based coding agents on cross-repository software change tasks.

Technical level: Advanced. The paper is readable without deep systems expertise, but interpreting it assumes familiarity with issue-resolution benchmarks, fail-to-pass / pass-to-pass test semantics, and agent scaffolds.

Scope: The paper introduces WideSWE, a benchmark of 120 real-world tasks that require a coding agent to implement one shared feature or bug fix across two or more repositories simultaneously, and reports an empirical study of seven agent configurations on it.

What This Paper Is About

Existing coding-agent benchmarks mostly measure whether an agent can resolve an issue inside a single repository. In real software ecosystems, however, a single feature or fix often spans several repositories — for example, adding the same capability to multiple language SDKs, or updating a library and every downstream project that depends on it. WideSWE asks whether coding agents can identify every repository that needs changing and deliver a coordinated, jointly correct set of changes across all of them.

Key Contributions

  1. A task formulation and benchmark for cross-repository coding. The authors formulate cross-repository completion as implementing one shared feature or bug fix across multiple scored target repositories, and release WideSWE: 120 real-world tasks covering 41 software ecosystems. The paper also reports that mining covered 103 active software ecosystems.

  2. Requirement-aligned test review. They identify mismatches between inherited PR tests and stated task requirements, then apply two review rules — relaxing implementation-specific constraints (private helper names, fixed file paths, exact error-message text) and removing checks for functionality that neither the issue, the PR, nor the prompt requested — while preserving required behavior and regression checks.

  3. An empirical study of cross-repository task completion. They evaluate seven agent configurations along dimensions including model, scaffold, task type (bug fix vs. feature), repository count, and language diversity, and pair a joint-versus-independent execution comparison with trajectory analysis.

  4. A released artifact. Instance metadata, base-commit identifiers, provenance records, prompts, hidden-test patches, environment definitions, scoring manifests, agent patches where licensing permits, and scripts for reproducing aggregate metrics are provided, with code at the linked GitHub repository.

Main Findings

  • Full task success is low and varies widely. Across seven agent configurations on 120 tasks, case-level success ranges from 10.83% to 42.50%. The best configuration, Codex CLI paired with GPT-5.6-sol, solves 42.50%; the other configurations achieve 10.83%–37.50%.

  • Repository-level progress outpaces task-level success. Codex CLI–GPT-5.6-sol solves at least one target repository in 83.33% of tasks, versus 42.50% full task success. Its overall repository success is 63.64%, F2P-complete 65.61%, and P2P-preserved 90.95%, at a mean of 86.7 API requests per task.

  • Incomplete functionality, not regressions, is the dominant failure. In 48 of the 69 failed Codex CLI–GPT-5.6-sol cases, at least one repository passes all its F2P tests while another does not. Failures due solely to P2P tests are uncommon, and P2P preservation rates stay high (84.16%–95.93% depending on configuration).

  • Three failure modes appear in the trajectories. The authors classify unresolved runs as (1) incomplete scope identification, (2) recognized work without delivery, and (3) post-edit failure. For Codex CLI–GPT-5.6-sol these account for 37.68%, 2.90%, and 59.42% of unresolved runs respectively; post-edit failure dominates for most configurations, reaching 73.33% for Qwen and 67.95% for Opus.

  • Models differ in where they break down. For Gemini 3.8 Flash, 52.34% of failures are recognized-work-without-delivery, and 85.98% of its unresolved tasks have no repository passing all F2P tests. Within that category, 82.14% of Gemini's runs remain in investigation, solution planning, or preparatory checks. DeepSeek (7.95%), Qwen (5.33%), and GLM (9.38%) show much lower rates of this mode, with omissions more often reflecting unrecognized scope.

  • Bug fixes and features behave differently. Qwen 3.8 Max outperforms Codex CLI–GPT-5.6-sol on bug fixes but falls behind on features (48.33% vs. 41.67% case success on bug fixes; 26.67% vs. 43.33% on features). All seven configurations score higher on multiple-language-family bug fixes than on single-family bug fixes, but feature tasks do not follow that pattern: Qwen reaches 37.14% on single-family features and 12.00% on multiple-family features, while GPT and Opus reach 28.00% on the multiple-family group.

  • More repositories mean lower task success without necessarily missing repositories. For Codex CLI–GPT-5.6-sol, repository success rises from 63.08% on two-repository tasks to 66.67% on three-repository tasks, while task success falls from 42.99% to 38.46%. Among its failed three-repository tasks, 87.50% modify all targets and 62.50% solve two of them.

  • The scaffold matters as much as the model. Holding GPT-5.6-sol fixed and switching from Codex CLI to Claude Code lowers task success from 42.50% to 32.50% while raising mean recorded API requests per task from 86.7 to 134.8.

  • Independent (per-repository) execution does not reliably help. On 89 tasks with identical prompts, joint execution solves 40.45% and independent execution solves 35.96%, with 22.47% of tasks changing outcome in either direction. Independent execution raises bug-fix success from 34.48% to 44.83% but lowers feature success from 43.33% to 31.67%. It uses 296.6 API requests per task versus 94.7 for joint execution — 3.13 times as many — without improving overall success.

  • A separate run helps most when nothing was attempted. Among repositories that fail in joint execution, 60.00% of those left unmodified succeed independently, compared with only 9.43% of those already modified.

  • Cross-repository context supports both implementation and verification. 13.48% of tasks and 9.95% of repositories succeed only under joint execution. Trajectory examples include an Ansible case where the joint run uses one repository's implementation to guide a compatible change in the other while the independent run makes incompatible assumptions, and a Symfony case where the joint run compares the two implementations, corrects a difference, and lets both pass while the independent runs leave errors in both.

  • Task composition. The 120 tasks comprise 60 bug fixes and 60 features, spanning 253 target repositories with 2,815 F2P and 22,139 P2P tests. Coordination patterns split into producer–consumer dependency (96 tasks, 80.0%) and parallel propagation (24 tasks, 20.0%). Workspaces contain up to 20 non-target context repositories, with a median of three; 27 tasks have none and 93 have at least one. Median reference-patch size is 13 files and 416 lines. Language-family labels split 86 single-family and 34 multiple-family tasks.

  • Selection funnel. From 1,729,171 PRs across 103 ecosystems, 109,233 records explicitly reference another repository in the same ecosystem, yielding 4,437 candidate groups after deduplication and checks; 2,188 had a latest PR merged on or after June 1, 2025; manual review retained 635; executable validation yielded 192 eligible cases (60 bug fixes, 132 features), from which all bug fixes and 60 features were kept.

Methodology in Plain English

The researchers started from the 200 GitHub organizations with the most repository stars (Gitstar Ranking, June 5, 2026) and identified 103 active ecosystems where several repositories support one product, platform, or technology. They collected pull requests merged since January 1, 2024 and looked for explicit links between repositories in the same ecosystem. Groups of linked changes that appeared to implement one shared request were filtered for deduplication, test availability, and Linux compatibility, then reviewed by hand to confirm that a single feature or fix truly required substantive changes in more than one repository. Surviving candidates were validated by execution: each target repository had to have at least one test that fails on the original code and passes after the reference change.

Prompts were assembled from the original issue and PR text, preserving wording and consolidating overlapping requirements, without adding implementation instructions. Hidden tests were taken from the reference patches and then manually revised under two rules aimed at removing implementation-specific constraints and unrequested functionality checks, after which the authors verified that the revised suites still reject the original code and accept the reference solutions.

Evaluation runs each agent on one prompt and a historical ecosystem workspace in which all repositories can be inspected and edited. A task counts as solved only if every target repository passes all of its fail-to-pass and pass-to-pass tests; context repositories are available but are not scored. Seven configurations were tested, each a scaffold-plus-model-plus-reasoning-effort pairing: Codex CLI (v0.147.0) with GPT-5.6-sol (high), and Claude Code (v2.1.139) with GPT-5.6-sol (high), Claude Opus 5 (high), Gemini 3.8 Flash (xhigh), DeepSeek V4 Pro (high), Qwen 3.8 Max (xhigh), and GLM 5.3 (xhigh). For the joint-versus-independent comparison, the same Codex CLI–GPT-5.6-sol setup was run on the 89 tasks whose prompts apply unchanged in both settings.

Why This Matters

Impact on research. Benchmarks such as SWE-bench, DeepSWE, and ProgramBench measure work inside one repository. WideSWE argues that this framing misses a common real-world unit of work — one requirement, several repositories — and shows that a state-of-the-art configuration solves only 42.50% of such tasks. It also shows that simply running an agent once per repository, which triples API usage, does not improve overall success.

Real-world applications.

  • Updating a shared behavior across a product's language SDKs, such as implementing the same capability in Go, Python, and Ruby repositories.
  • Propagating a new capability or interface through a library and the downstream repositories that depend on it, as in the Sentry PHP SDK case that requires updates to two further repositories.
  • Keeping parallel implementations consistent, as in the Symfony case where two repositories must implement identical behavior and one serves as a behavioral reference for the other.
  • Platform and framework changes that must land jointly, such as edits spanning Kubernetes repositories or the Godot engine and a native SDK.

Industry relevance. The failures observed are the ones that matter to maintainers: silently narrowing scope to one package, recognizing downstream work and never delivering it, and landing changes that look complete but violate requested behavior. These are precisely the risks in monorepo-adjacent and multi-package release workflows, where a partial change can ship a broken interface to consumers.

Future Directions

  • Reducing post-edit failure, the most common breakdown, by making agents verify that their changes jointly satisfy the request rather than stopping when one repository's tests pass.
  • Closing the gap between scope identification and delivery, given that delivery-omission rates exceed 50% for Gemini and that runs like the DeepSeek Prettier/yaml-unist-parser case and the Gemini Graphon/Dify case spend hundreds of API calls without finishing recognized work.
  • Deciding when to split work rather than run it jointly: independent execution helped bug fixes (34.48% to 44.83%) but hurt features (43.33% to 31.67%), so a policy for choosing between settings is an open question.
  • Improving coordination as repository count grows, since task success fell from 42.99% to 38.46% between two- and three-repository tasks even as repository-level success rose.
  • Determining whether the three observed failure categories are stable across future models and scaffolds, and whether the two requirement-alignment test-review rules generalize to benchmarks beyond WideSWE.

Target Audience

Researchers and practitioners working on coding agents, software-engineering benchmarks, and LLM evaluation; maintainers and release engineers responsible for multi-repository ecosystems and language SDKs; and tool builders designing scaffolds, permissions, and execution strategies for autonomous repository work. Readers primarily interested in single-file code generation will find less of direct relevance.

Authors’ abstract

Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.

Read the original paper