Research
OSS-CRS: Liberating AIxCC Cyber Reasoning Systems for Real-World Open-Source Security
Overview Research area: Computer security and software engineering, specifically automated vulnerability discovery and repair (cyber reasoning systems), and the systems infrastructure needed to deploy

- arXiv
- 2603.08566
- Published
- 2026-03-09
- Authors
- Andrew Chin, Dongkwan Kim, Yu-Fu Fu, Fabian Fleischer, Youngjoon Kim, HyungSeok Han, Cen Zhang, Brian Junekyu Lee, Hanqing Zhao, Taesoo Kim
AI summary
Overview
- Research area: Computer security and software engineering, specifically automated vulnerability discovery and repair (cyber reasoning systems), and the systems infrastructure needed to deploy such systems on real open-source software.
- Technical level: Advanced. The paper assumes familiarity with fuzzing, sanitizers, Docker/Kubernetes, LLM-based agents, and CI/CD security workflows.
- Scope in one sentence: The paper presents OSS-CRS, a locally deployable, budget-aware framework that decouples the seven open-sourced AIxCC cyber reasoning systems from the now-defunct competition cloud, and demonstrates the ported first-place system (Atlantis) finding 10 previously unknown bugs across 8 OSS-Fuzz projects.
What This Paper Is About
DARPA's AI Cyber Challenge (AIxCC, 2023–2025) produced seven autonomous cyber reasoning systems (CRSs) that could not only find vulnerabilities but also confirm them with a proof of vulnerability (PoV) and synthesize a validated patch — and all seven teams open-sourced their systems. However, more than half a year later, those systems remain largely unusable outside their original teams because each is bound to competition-specific cloud infrastructure (Microsoft Azure plus Kubernetes) that has since been decommissioned. This paper diagnoses why that adoption gap exists and builds OSS-CRS, an open framework that lets these CRSs run and be combined on real-world open-source projects on ordinary local hardware, with enforcement of CPU, memory, network, and LLM spending limits.
Key Contributions
- An empirical analysis of all seven AIxCC finalist CRS codebases, identifying three recurring deployment barriers: infrastructure duplication, cloud lock-in, and monolithic design.
- OSS-CRS, an open-source framework that addresses those barriers through a unified three-phase execution model (prepare, build-target, run) with a standard interface (libCRS), budget-aware resource management across CPU, memory, and LLM usage, and support for combining CRS techniques across independently developed systems.
- Real-world validation by porting Atlantis, the first-place AIxCC system, which required over 20 Azure VMs, 42 TiB of cloud storage, and runtime Azure SDK calls in its original form; the port discovered 10 previously unknown bugs, three of high severity, across 8 OSS-Fuzz projects.
- Five integrated CRSs and public release — crs-libfuzzer, Atlantis-C, Atlantis-Java, Atlantis-MultiLang, and ClaudeCode — spanning traditional fuzzing through LLM-powered multi-language analysis and LLM-based patch generation, all made publicly available.
Main Findings
-
Three deployment barriers block reuse: infrastructure duplication (each of the seven teams independently rebuilt the same platform services), cloud lock-in (every system targets the competition's Azure and Kubernetes environment, which was shut down), and monolithic design (analysis techniques are embedded with no modular internal interfaces).
-
Duplication is concrete and measurable: across the seven finalists, container counts range from 1 (RoboDuck) to 53 (Artiphishell), with Atlantis at 9+, Buttercup at 14+, BugBuster at 16, and Lacrosse at 3+. Teams converged on similar platform roles using different tools — Terraform with Kubernetes or Helm, and coordination backends such as Kafka, RabbitMQ, Redis, and PostgreSQL. Five of seven teams deployed LiteLLM as their LLM gateway, and each independently built its own LLM budget tracking, cost enforcement, and model routing logic on top of it.
-
Cloud dependency is severe: Atlantis requires over 20 Azure VMs, 42 TiB of cloud storage, and runtime Azure SDK calls to scale Kubernetes node pools. Only half the teams (Buttercup, RoboDuck, FuzzingBrain, Artiphishell) added any local execution support through standalone releases or local deployment guides.
-
No CRS is composable: Table I confirms that none of the seven finalists exposes interfaces for component-level extraction. A researcher cannot, for example, combine Atlantis's fuzzer with Buttercup's patcher without reimplementing one inside the other. Even Zhang et al.'s survey of the seven finalists could describe each team's methods but not compare them experimentally.
-
Budget control is a first-class concern: AIxCC allocated $50,000 in LLM credits per team, and without budget controls a single CRS run can exceed $1,000 per hour.
-
The framework works at real scale: OSS-CRS integrates five CRSs. The simplest, crs-libfuzzer, wraps libFuzzer with a two-line build script; the most complex, the three ported Atlantis sub-systems, retain their original bug-finding logic while dropping competition-specific infrastructure.
-
Porting is tractable: Atlantis-MultiLang took approximately 3 person-days to integrate, Atlantis-C took over 4 person-days (it was more tightly coupled to competition build pipelines and container orchestration), and Atlantis-Java and ClaudeCode took approximately 1 person-day each. Only configuration, orchestration logic, and artifact I/O boundaries changed — not the core analysis algorithms.
-
Zero-day results: Running against 8 OSS-Fuzz projects ranging from 3 kLoC to 1,960 kLoC, spanning C/C++ and Java and domains including databases, parsers, network servers, and cryptographic libraries, OSS-CRS produced 10 previously unknown bugs, three of high severity. Three have been fixed by upstream maintainers, one is confirmed but not yet patched, and six are pending initial response. The bugs are mostly memory-safety issues in C, but the set also includes logic bugs and undefined-behavior flaws classed as CWE-476, CWE-674, and CWE-681 — null-pointer, schema-validation, and numeric-conversion problems rather than only fuzzer-class crashes.
-
Experiment setup: All runs used a single machine with 32 CPU cores and 128 GB RAM on Ubuntu 22.04 with Docker 27. Each campaign allocated 16 cores and 64 GB RAM to the bug-finding CRS with a 24-hour timeout per target, and LLM budgets of $50 per campaign using a mix of Claude and GPT-4o.
-
One component was deliberately excluded: Atlantis's concolic hybrid fuzzing component was not ported because its instrumentation introduced compatibility issues across diverse target projects; in the original competition it contributed only 1.7% of Atlantis's results.
Methodology in Plain English
The authors started by reading the released code and deployment artifacts of all seven AIxCC finalists to work out, concretely, why nobody outside the original teams could run them. That analysis produced three named barriers, which then became the design requirements for a new platform.
Rather than build another CRS, they built the platform underneath one. OSS-CRS splits the workflow into three separate phases. The prepare phase builds the CRS's own container images, which have no knowledge of any target project, so the setup cost is paid once and reused. The build-target phase compiles the specific target project using OSS-Fuzz's official build flows, so the framework inherits the build environment that project maintainers already support across heterogeneous build systems like Make, CMake, Autoconf, Bazel, and Meson. The run phase launches the CRS containers together on isolated Docker networks and executes the campaign.
CRSs talk to the platform through a small Python library called libCRS, which is injected into every container. It handles publishing build outputs, sharing files within a multi-container CRS, submitting and fetching artifacts such as PoVs, seeds, and patches, and running the patch validation loop. Because that validation loop — apply the patch, rebuild incrementally, re-run the PoV, run the regression tests — is handled by a shared "builder sidecar," a CRS like ClaudeCode contains no build logic at all and can focus purely on generating patches.
Resources are managed explicitly. Each CRS declares a cpuset and a memory limit enforced through Docker cgroups, gets its own Docker network, and receives a US-dollar LLM budget. The LiteLLM-based proxy issues a unique API key per CRS, tracks cumulative cost, and rejects requests once the budget is exhausted. CRSs exchange artifacts only through a filesystem-based exchange directory, with content-hash filenames so duplicate discoveries appear once; they never communicate directly, which means one crashing CRS does not take down the others.
For validation, the authors ported three of Atlantis's bug-finding sub-systems as independent CRSs, cut the system's three LiteLLM proxies down to one, replaced Kubernetes orchestration with flat Docker, replaced the competition scoring API with libCRS submission, and then ran the result against eight OSS-Fuzz projects on a single machine. Patches produced were automatically validated by the libCRS pipeline and then manually reviewed before being sent to upstream maintainers through each project's preferred disclosure process, including GitHub Security Advisories and direct email.
Why This Matters
Impact on research. The paper argues that the bottleneck in autonomous vulnerability remediation is not technique capability but robust integration — a conclusion it attributes to Zhang et al.'s systematization of the seven finalists. By turning monolithic competition systems into composable, locally runnable components, OSS-CRS makes controlled cross-CRS experiments and ablation studies possible for the first time. A researcher can now mix the strongest component from one team with the strongest from another, or run multiple CRSs on the same source tree with matched inputs so results are directly comparable and reproducible.
Real-world applications.
- CI/CD security gates: A pipeline submits a pull-request diff and receives either a PoV or confirmation of no bugs through a stable machine-readable interface, checking only the changed code region rather than the whole codebase.
- Turning static analysis into fixes: Security practitioners can feed SARIF reports from existing static analyzers into the system and get back validated patches, with spending limits enforced automatically.
- Maintainer relief on bug reports: The curl project shut down its bug bounty program after AI-written submissions overwhelmed reviewers with unconfirmed findings, and FFmpeg maintainers criticized Google for reporting valid bugs without patches, calling them "CVE slop." OSS-CRS targets exactly this gap by producing patches validated against the triggering PoV and the project's regression tests.
- Budget-constrained deployments: Organizations that cannot afford uncontrolled LLM costs can cap spending per CRS in dollars and compare efficiency fairly — a CRS achieving the same result within $50 is more efficient than one requiring $500.
Industry relevance. Competition-grade autonomous security reasoning was previously usable only by the teams that built it, on cloud infrastructure that has since been decommissioned. OSS-CRS shows that capability can be decoupled from its infrastructure and run on a single commodity machine, which lowers the barrier for open-source maintainers, security teams, and tool vendors. The framework's adoption of the OSS-Fuzz project format means any integrated CRS can target over 1,000 OSS-Fuzz projects without per-project customization.
Future Directions
- Port the remaining finalist CRSs. The paper validates OSS-CRS with only one AIxCC system (Atlantis) and states the authors are working on porting other finalists; additional ports may reveal interface gaps or design assumptions that a single system does not exercise.
- Quantify ensemble gains. The paper demonstrates that cross-CRS artifact flow works in practice but explicitly states that quantifying performance gains over single-CRS runs remains future work.
- Smarter artifact deduplication. The current content-hash deduplication in the exchange sidecar is described as a baseline; the authors propose extending it with coverage-based deduplication that keeps only seeds increasing overall coverage, or stack-trace-based deduplication that groups PoVs by crash signature to reduce redundant triage.
- Extend beyond OSS-Fuzz. The framework currently requires projects compatible with OSS-Fuzz — a standardized build script, containerized environment, language restrictions, and at least one fuzz harness. Projects not yet onboarded require harness development, which the paper places outside the framework's current scope.
Target Audience
This paper is most valuable to CRS developers and security tooling engineers who want a stable integration contract so they can focus on analysis logic instead of platform engineering; to security researchers who need a consistent environment for comparing and combining vulnerability discovery and patching techniques across systems; to security practitioners and platform teams evaluating deployable, budget-aware bug-finding and bug-fixing workflows; and to maintainers of large open-source projects who are the intended beneficiaries of validated PoVs and patches. Readers without background in fuzzing, sanitizers, or container orchestration will find the systems sections dense, though the motivation and barrier analysis in the early sections is broadly accessible.
Authors’ abstract
DARPA's AI Cyber Challenge (AIxCC) showed that cyber reasoning systems (CRSs) can go beyond vulnerability discovery to autonomously confirm and patch bugs: seven teams built such systems and open-sourced them after the competition. Yet all seven open-sourced CRSs remain largely unusable outside their original teams, each bound to the competition cloud infrastructure that no longer exists. We present OSS-CRS, an open, locally deployable framework for running and combining CRS techniques against real-world open-source projects, with budget-aware resource management. We ported the first-place system (Atlantis) and discovered 10 previously unknown bugs (three of high severity) across 8 OSS-Fuzz projects. OSS-CRS is publicly available.