Research
Who Does Withholding Delay? A Game-Theoretic Model of Open-Weight AI Release Under Asymmetric Proliferation
Overview Research area: AI safety and governance, specifically the decision theory of open-weight model release, formalized with Stackelberg game theory and cybersecurity economics. The paper sits at
- arXiv
- 2607.22957
- Published
- 2026-07-24
- Authors
- Daniel Commey
AI summary
Overview
Research area: AI safety and governance, specifically the decision theory of open-weight model release, formalized with Stackelberg game theory and cybersecurity economics. The paper sits at the intersection of AI policy, industrial organization, and security economics.
Technical level: Advanced. The paper states formal propositions with proofs, uses exponential hazard models, discounted-welfare integrals, and a numerical policy-comparison implementation, though a summarized "practitioner decision rule" is written for non-specialists.
Scope: The paper builds a three-stage Stackelberg model in which a laboratory choosing among four release policies must account for actor-specific substitution times, and derives conditions under which withholding a dual-use model actually delays harmful actors more than defenders.
What This Paper Is About
Restricting access to a dual-use AI model only counts as precautionary if it slows harmful actors more than it slows defenders, and that condition differs by actor: a state agency or organized criminal group may obtain a substitute through theft, distillation, intermediated API access, independent development, or a foreign release, while a small utility or open-source maintainer may have no comparable route. The paper asks "which actors does withholding actually delay?" and models a laboratory choosing among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights. The goal is to identify the specific quantities a release review would need to measure — actor-specific substitution times, marginal capability gains, deployment rates, defensive reach, newly enabled misuse, and nonrecallable losses — before treating restriction as precautionary.
Key Contributions
-
Formalizes access inversion. Proposition 1 gives a closed-form benchmark, $A_i = \lambda_i / [\rho(\lambda_i + \rho)]$, showing that restriction creates a positive discounted adversary access advantage if and only if the sophisticated-adversary acquisition rate exceeds the defender rate ($\lambda_S > \lambda_D$).
-
Formalizes asymmetric empowerment at a finite horizon. Proposition 2 shows the marginal empowerment of immediate release to actor $i$ over a horizon $H$ is $\Delta C_i(H) = q_i e^{-\lambda_i H}$, so with equal usefulness the actors least able to obtain a substitute gain the most capability.
-
Derives a proliferation-reversal threshold. Proposition 3 shows that in a linear damage benchmark the welfare difference between broad release and control is strictly increasing in $\lambda_S$, yielding a unique threshold $\lambda_S^* = \rho,\theta/(1-\theta)$ with $\theta = -\rho,\Psi(0)/(\alpha q_S)$ above which broad release overtakes control.
-
Adds a numerical implementation and an evidence dataset. A nonlinear model produces a nonempty policy region for each release tier, three nested 2,048-point deterministic designs show how policy shares change with the parameter bounds, and a separate grid examines actor-specific deployment delays after release.
Main Findings
-
Access inversion (Proposition 1). Under independent exponential substitute-acquisition times, discounted access exposure is $A_i = \lambda_i/[\rho(\lambda_i+\rho)]$, and the adversary–defender gap is $\frac{1}{\rho}[\frac{\lambda_S}{\lambda_S+\rho} - \frac{\lambda_D}{\lambda_D+\rho}]$. Restriction creates a positive discounted adversary access advantage if and only if $\lambda_S > \lambda_D$.
-
Asymmetric empowerment (Proposition 2). Immediate release changes access probability from $a_i(H) = 1 - e^{-\lambda_i H}$ to one, so $\Delta C_i(H) = q_i e^{-\lambda_i H}$. With $q_S = q_D = q$ and $\lambda_S > \lambda_D$, $\Delta C_D(H) > \Delta C_S(H)$ for every $H > 0$; more generally the condition is $\frac{q_D}{q_S} > e^{-(\lambda_S - \lambda_D)H}$.
-
Neither effect makes openness optimal on its own. The policy ranking also depends on opportunistic misuse ($m_r$), the offense–defense conversion ratio $\omega = \alpha/\beta$, defensive spillovers through reach $n_r$ and network productivity $\eta$, safeguard friction, and nonrecallable losses ($I_r$).
-
A proliferation-reversal threshold exists. $\Psi(\lambda_S) = W(B) - W(\mathsf{C})$ is strictly increasing with $\frac{\partial \Psi}{\partial \lambda_S} = \frac{\alpha q_S}{(\lambda_S+\rho)^2} > 0$. If $\Psi(0) < 0 < \lim_{\lambda_S \to \infty}\Psi(\lambda_S)$, there is a unique $\lambda_S^$ with controlled access preferred below it and broad release preferred above it, given by $\lambda_S^ = \rho\frac{\theta}{1-\theta}$ where $\theta = -\frac{\rho \Psi(0)}{\alpha q_S} \in (0,1)$.
-
A defender-first window has narrow conditions. Selected defenders receive the model at time zero but protection exists only at $T_P \sim \text{Exp}(\mu)$, so a useful private window requires $T_P < \min{T_S, \tau}$. The benchmark omits four failure modes: incomplete recipient coverage, invalid mitigations, slow rollout, and leakage during the window.
-
Removable safeguards are not useless. Safeguarded open weights let a sophisticated user strip the defaults, but the defaults can still deter enough opportunistic misuse for the tier to be reasonable.
-
A state-capability counterexample. When a state already holds an adequate substitute, the welfare difference between safeguarded open release and control is $W(\mathsf{O_g}) - W(\mathsf{C}) = \frac{\Delta b + d - m}{\rho} - \Delta I$, so safeguarded open release is better when $\Delta b + d - m > \rho\Delta I$; continued control is justified only when avoided misuse and tail costs outweigh what defenders lose.
-
Baseline calibration. A defender rate $\lambda_D = 0.70$ corresponds to an expected effective-acquisition time of about 1.43 years; $\lambda_S = 0.95$ corresponds to about 1.05 years.
-
Effort microfoundation yields a unique equilibrium. With $\lambda_i(e_i) = \bar{\lambda}_i + k_i e_i$ and quadratic effort cost $c_i e_i^2/2$, payoffs are strictly concave and separable, so best responses form a unique Nash equilibrium satisfying $c_i e_i^* = \frac{v_i k_i}{[\bar{\lambda}_i + k_i e_i^* + \rho]^2}$. Higher access value, higher acquisition productivity, and lower effort cost each raise the equilibrium rate, so $\lambda_S > \lambda_D$ can arise endogenously.
-
Competing channels aggregate by minimum, not sum of times. With independent channels, $T_i = \min_k T_{ik} \sim \text{Exp}(\sum_k \lambda_{ik})$; a channel needing artifact, compute, and integration in sequence delivers access at the maximum of those times.
-
Observed release paths are not the same as observed access. GPT-2's largest weights took 264 days to arrive; Llama 3.1, DeepSeek-R1, and gpt-oss shipped weights on announcement day; Moonshot's Kimi K3 (2.8 trillion parameters, one-million-token context) opened hosted access immediately with weights scheduled for July 27, past the paper's July 24 source cutoff, and its deployment guidance recommends a supernode of at least 64 accelerators.
-
The July 2026 Hugging Face incident documents an access asymmetry. Hugging Face's response team worked through more than 17,000 recorded attacker events; its first attempt used frontier models behind commercial APIs that refused requests containing real attack commands, exploit payloads, and command-and-control artifacts, and the analysis was completed on a self-hosted GLM 5.2. OpenAI separately attributed the intrusion to GPT-5.6 Sol and a more capable pre-release model operating with reduced cyber refusals inside an internal capability evaluation.
-
Numerical results confirm a nonempty policy region per tier. The nonlinear implementation produces a nonempty policy region for each release tier, and three nested 2,048-point deterministic designs show how policy shares change with the parameter bounds.
-
The stated decision rule is conditional, not permissive. Control is welfare-improving only when it materially delays high-harm actors; a defender-first window only when defenders can ship before adversaries catch up; safeguarded release once substitutes have eroded the control advantage and defaults still deter enough misuse; minimally restricted release requires the stronger finding that safeguard friction on legitimate work exceeds whatever deterrence the defaults provide.
Methodology in Plain English
The paper sets up a three-stage Stackelberg game. A laboratory commits first to one of four release policies: controlled hosted or gated access, a defender-first window ($\mathsf{P}(\tau)$) that gives selected recipients an early copy before public release at time $\tau$, immediate open-weight release with default safeguards, or immediate minimally restricted open weights.
Three actor classes sit downstream: sophisticated adversaries that can chase substitutes even under restriction, opportunistic adversaries that mostly appear when public access is frictionless, and a distributed defender population that produces protection through both direct model use and the heterogeneous participation of many independent maintainers.
Under restriction, sophisticated adversaries and defenders each spend costly effort trying to obtain an adequate substitute. Substitute-acquisition times are modeled as independent exponential random variables with rates $\lambda_S$ and $\lambda_D$ — the constant-hazard benchmark — and the paper shows how multiple independent acquisition channels combine through competing risks. The laboratory anticipates these responses because follower payoffs are separable.
The analyst then compares policies using a social welfare function: an expected discounted flow of broad benefit minus expected harm, minus a one-time, policy-specific irreversibility cost for losses that cannot be recalled once weights are copied. Harm rises with the strategic capability gap between adversaries and defenders and falls with defensive reach, and includes a separate reduced-form term for opportunistic misuse. Four assumptions frame the analysis: the laboratory can commit to its announced tier and timing; restricted-access acquisition is stationary and independent across actors; access is binary and an acquired substitute is adequate for the capability under study; and welfare is an expected discounted flow plus a one-time irreversibility term.
Analytic propositions handle the linear benchmark. A nonlinear numerical implementation relaxes linearity, compares all four policies, and stress-tests the ranking over three nested parameter boxes, with a separate grid for actor-specific deployment delays after release ($X_{i,B}(t) = \mathbf{1}{t \geq D_{iB}}$). A short evidence dataset records observable release timing alongside the quantities that remain out of reach, drawn from recent release, cyber-evaluation, and incident-response cases.
Why This Matters
The paper reframes a release decision that is usually argued in terms of model capability into a question about actor-specific delay: for each population that matters, estimate the capability release would add over a decision horizon and the time each actor needs to find a substitute without you. That is a measurable question rather than a categorical one, and it makes the normative claim conditional on data an organization can in principle collect.
Impact on research. The paper positions itself against two neighboring literatures. Prior game-theoretic treatments of openness — such as Mladenovic, Courville, and Gidel, whose players are still building the model when they choose, and Landolt et al.'s bilevel Stackelberg study of defender-first sequencing — do not make the value of sequencing depend on adversary substitution, defender coverage, and deployment. This paper additionally formalizes one piece of the marginal risk identified by Kapoor et al., who argue that open-model risk must be measured against substitutes.
Real-world applications:
- Release review inside model developers. The paper argues a review should estimate actor-specific substitution times, marginal capability gains, deployment rates, defensive reach, newly enabled misuse, and nonrecallable losses before treating restriction as precautionary.
- Enterprise and public-sector deployment choices. The paper contrasts monitored API access (updates and provider-side controls) against open weights (local data residency, deep customization, predictable availability, freedom from a vendor's next unilateral change), citing Meta's report that Llama 3.1 405B can be run on a developer's own infrastructure at roughly half the inference cost of GPT-4o, and OpenAI's 2025 gpt-oss release running the larger variant on a single 80-GB GPU and the smaller in 16 GB of memory.
- Incident response planning. The Hugging Face case motivates the recommendation to retain safeguards on hosted systems while maintaining a vetted self-hosted fallback for workflows that hosted models reject.
- National policy and export-control design. The paper treats national policy as an input into $\lambda_i$, $n_r$, $f_r$, and the newly enabled misuse term, to be re-checked whenever controls, releases, or deployment costs move — noting that U.S. chip controls may slow training of future substitutes but cannot prevent a foreign lab from publishing weights already trained, and Chinese service rules govern monitored domestic deployment but confer no recall power over copies held abroad.
Industry relevance. The paper's release-record proposal is directly operational: track four milestones per model (announcement, hosted availability, weight availability, and actor-specific deployment), and update the record whenever a substitute release or a cost drop moves the picture. It also explicitly notes that even weight publication overstates effective access — what a small hospital or a large national lab can actually run depends on memory, accelerator count, integration work, and which smaller derivatives exist.
Future Directions
-
Measure the substitution rates directly. The paper states that empirical work would target $\lambda_S$ and $\lambda_D$ directly rather than the effort primitives underneath them, and that national policy should be re-checked whenever controls, releases, or deployment costs move.
-
Move beyond binary access. The maintained assumptions rule out partial or task-specific substitutes, which the paper says would need a multi-state model; a model can substitute for vulnerability discovery without substituting for autonomous intrusion, incident reconstruction, or patch deployment.
-
Model coupled contests and common shocks. The baseline assumes followers' acquisition efforts are independent and rules out common shocks; a coupled contest would add cross-effects between the followers' efforts.
-
Build out the release-record dataset. The five examined releases are described as far too few to estimate population frequencies or acquisition rates, so the paper calls for release records tracking announcement, hosted availability, weight availability, and actor-specific deployment per model.
-
Quantify nonrecallable loss from first principles. The model ranks policies under stated tail-cost assumptions rather than deriving the social value of catastrophic outcomes, leaving $I_r$ as an input.
Target Audience
The paper is written for two audiences at once. Its primary reader is a researcher or analyst working on AI governance, AI safety policy, or the economics of dual-use technology who wants a formal account of when restriction actually delays harmful actors. Its secondary reader is a practitioner — someone preparing or reviewing a release decision, an enterprise choosing between hosted API access and downloadable weights, or a policy staffer translating export-control and diffusion strategy into concrete parameters. The paper explicitly offers a route for operational readers to jump straight from the practitioner decision rule to the industry cases and the decision framework without working through the model and proofs.
Note: the provided content is truncated mid-proof of Proposition 3, so the full statements of Proposition 4, Proposition 5, and the numerical results are not available in this excerpt and are therefore not reported here.
Authors’ abstract
Restricting access to a dual-use AI model is precautionary only if it delays harmful actors more than defenders. That condition varies across actors: a state agency or organized criminal group may obtain a substitute through theft, distillation, intermediated access, independent development, or a foreign release, while a small utility or open-source maintainer may have no comparable route. We model a laboratory choosing among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights. Access inversion occurs when restriction gives an access advantage to adversaries that obtain effective substitutes faster than defenders. Asymmetric empowerment occurs when immediate release adds the most capability to populations least likely to possess a substitute. The policy ranking also depends on relative usefulness, opportunistic misuse, offense-defense conversion, defensive spillovers, safeguard friction, and nonrecallable losses. A linear benchmark yields a unique adversary-substitution threshold above which broad release overtakes control when the endpoint conditions hold. A defender-first window has value when selected defenders deploy protection before adversaries catch up, and removable safeguards remain useful when they deter enough opportunistic misuse. A nonlinear implementation gives each release tier a nonempty policy region. Three nested 2,048-point deterministic designs assess sensitivity to parameter bounds, and a separate grid examines actor-specific deployment delays after release. Release, cyber-evaluation, and incident-response cases identify the quantities a release review should estimate: actor-specific substitution times, marginal capability gains, deployment rates, defensive reach, newly enabled misuse, and nonrecallable losses.