Research
Exploring Systems-Thinking Approaches to Loss of Control Risk
Overview Research area: AI Safety & Ethics — specifically Loss of Control (LoC) risk in frontier AI, analysed through systems-safety and systems-theoretic methods rather than model-level evaluations.

- arXiv
- 2606.13474
- Published
- 2026-06-11
- Authors
- Aurelio Carlucci, Sean P. Fillingham, James Walpole, Jakub Kryś
AI summary
Overview
Research area: AI Safety & Ethics — specifically Loss of Control (LoC) risk in frontier AI, analysed through systems-safety and systems-theoretic methods rather than model-level evaluations.
Technical level: Intermediate. The paper is conceptual and qualitative; no quantitative modelling or code is presented, but it assumes familiarity with risk-analysis vocabulary (control loops, hazards, unsafe control actions, resonance) and with the frontier-AI safety-framework landscape.
Scope: The paper asks whether established systems-safety methods (STECA, STPA, and FRAM) can surface internal-deployment Loss of Control risks in a generic frontier-lab coding-agent scenario that model-level evaluations would miss.
What This Paper Is About
The authors define internal-deployment Loss of Control as a state in which human controllers can no longer reliably constrain, audit, reverse, or halt AI-mediated changes to internal code, infrastructure, evaluation, or deployment processes within the time window required to prevent serious organisational or societal harms. They observe that most existing LoC frameworks focus on model-level misalignment (deception, scheming, behavioural drift), while the actual deployment setting is a sociotechnical system spanning developers, integration pipelines, monitoring systems, and organisational policies. The goal is not a definitive risk assessment but a demonstration that systems-theoretic methods apply productively to this setting and to specify exactly which disclosures would be needed to carry the analysis further.
Key Contributions
- An operational framing of internal-deployment LoC that locates the risk in the inability to constrain, audit, reverse, or halt AI-mediated changes to code, infrastructure, evaluation, or deployment processes in time to prevent serious harm — a definition aimed at the deployment system rather than the model.
- A proof-of-concept application of three complementary systems-safety methods (STECA, STPA, FRAM) to a generic frontier-lab coding-agent scenario reconstructed from public materials, showing what each method surfaces and where each reaches its limits.
- A set of concrete governance and operational findings, including unspecified controller identities in published frameworks, a missing policy-stewardship loop upstream of enforcement mechanisms, and a permissive/restrictive asymmetry in deployment authority.
- A regulatory proposal distinguishing structural from operational transparency requirements, including a two-tier disclosure regime (aggregate or redacted public indicators plus full operational data to a designated auditor under confidentiality).
Main Findings
-
Published frameworks leave governance responsibilities externally unverifiable. STECA Pass 1 yields about 50 distinct completeness and consistency items across a 21-row control-structure model, following Table 7 of Fleming (2015). The published record is dense at the enforcement and monitoring layer (access controls, sandboxing, operational parameter checks such as compute utilisation, bandwidth, and filesystem access, the multi-stage monitoring pipeline, and the CI/CD code-review stack) but markedly sparser at policy stewardship and coordination layers.
-
Four control actions are stated without a named executor. Access policy stewardship, catastrophic risk assessment and quarterly summaries, pause/restrict/reroute authority, and incident reporting each constitute a whole control loop whose controller identity, feedback channel, and process model are unspecified in public records. The authors note that the two items corresponding to enacted regulatory obligations are therefore expected to be externally verifiable.
-
Enforcement controllers lack an upstream policy steward. Four enforcement controllers are named and specified (access control, command allowlisting, automated invariant tests, and sandboxing), but the controller that authors the policy each one enforces is not specified. Without a stewardship loop there is no feedback on cumulative policy drift, since the code-review stack evaluates individual changes rather than the cumulative posture they produce.
-
A permissive/restrictive asymmetry exists in deployment authority. One named controller is authorised to expand deployment scope (temporary pre-audit deployment), but no structural counterpart is authorised to contract it.
-
Quantitative adequacy gaps are the bulk of enforcement-layer completeness findings. Named examples include the 30-minute monitoring latency against action velocity, human review bandwidth against codebase change volume, and coverage exclusions on specific code categories.
-
STPA surfaces latency-driven loss scenarios rather than component failures. Hazard H-4 is that safeguards do not match the actual capabilities or autonomy level of the deployed model [L-1, L-6]. The highlighted unsafe control action UCA-4.3a is that Safety and Security Oversight (SSO) provides the Kill Instance command too late, after the AI system has already produced dangerous or harmful outputs. Two generalised loss scenarios produce this: GLS-3 (Procedural Delay, where the organisational decision protocol introduces latency) and GLS-11 (Feedback Missing or Delayed, where the signal is queued or lost before reaching the controller).
-
STPA's control structure comprises three controllers. Safety and Security Oversight (SSO), Company Software Engineer (SWE), and Frontier AI Coding Agent, with the Company Codebase as the primary controlled process. SWE and the AI Agent are both controllers and controlled processes; eight system hazards are identified in total.
-
FRAM identifies degradation of a well-designed structure through ordinary operational variability. Six functions are modelled — four foreground (F1 develop software using an AI agent; F2 generate and execute code changes; F3 integrate changes into the codebase; F4 monitor and oversee agent activity) and two background (B1 provide development infrastructure and compute; B2 set and maintain policies, permissions, and constraints).
-
Four resonance patterns are derived. Pattern 1, Correlated control drift, where permission exceptions accumulate, non-firing monitoring rules are removed, review requirements stay static while PR volume grows, and the system prompt is not updated after a model upgrade — all pushing in the same direction of reduced control margin [H-4, H-6]. Pattern 2, Model swap discontinuity, where a model upgrade changes F2's variability overnight while F4 remains calibrated to the previous model, creating an instant rather than gradual gap [H-4, H-5]. Pattern 3, Codebase saturation, where a feedback cycle (F2 produces diffs and rationale, F1 submits a PR, F3 integrates into the codebase, F2 reads the codebase as context) causes agent-authored content to become the baseline for distinguishing normal from abnormal [H-7]. Pattern 4, Rationale-anchored review, where the agent's explanation is bundled with the artefact being assessed and anchors reviewer expectations even for careful reviewers, reducing backstop independence [H-2, H-3].
-
Some FRAM patterns were discarded. An exfiltration-via-side-effects pattern proved unrealistic under the baseline controls, and a control-infrastructure-degradation pattern reduced to standard software quality management.
-
The methods are complementary and the authors did not prioritise findings. STECA identifies what the public record leaves unspecified and where STPA must make assumptions; STPA identifies hazards and unsafe control actions in the designed structure; FRAM shows how those hazards could emerge through ordinary variability. The authors deliberately decline severity and likelihood estimates, arguing that operational frequency data and consequence models are not publicly available for frontier AI deployments and that producing such estimates from the public record alone would give a false impression of quantitative grounding.
Methodology in Plain English
The authors define a generic frontier AI developer as a sociotechnical system, with internal deployment of frontier AI models as the controlled process and the organisational boundary as the system boundary. Because no named company's internals are observable, they reconstruct a representative scenario from public disclosures by frontier AI developers, enacted regulation, and recent AI Control work, and they assume a baseline set of controls broadly consistent with public materials from Anthropic, OpenAI, and that literature (model-weight and sensitive-resource security; substantial logging and automated monitoring of internal agent use; manual review before agent-generated code is integrated; sandboxing and restricted execution; and reporting, governance, and incident response channels).
Three methods are then applied at a deliberately high level of abstraction. STECA treats published safety frameworks, system cards, risk reports, and regulation as a Concept-of-Operations document — the authors author a ConOps for a generic developer — and runs two passes: Pass 1 builds the control-structure model from published framework and regulatory material only, registering explicit absences rather than filling gaps, while Pass 2 evaluates the derived requirements against evidential material such as operational research, whistleblower accounts, and investigative disclosures, assigning each an evidenced, violated, or unknown verdict, with every unknown becoming a specific disclosure requirement. STPA functionally models the system as a hierarchy of controllers imposing constraints on controlled processes and enumerates unsafe control actions and generalised loss scenarios. FRAM models the system as coupled functions whose ordinary output variability can combine to produce emergent outcomes without any local failure.
The authors are explicit that this is one defensible application of each method, not the unique output, because identifying losses, drawing boundaries, and clustering loss scenarios are interpretive acts; they note that the methods were run in parallel for practical reasons, though they suggest a natural ordering of STECA, then STPA, then FRAM.
Why This Matters
Impact on research. The paper argues that systems-theoretic methods can be productively applied to frontier AI, and that the resulting findings — sociotechnical control structures, governance feedback loops, and operational variability — would not be surfaced by capability evaluations, alignment benchmarks, or propensity testing. It also flags that no established risk-management standard for general-purpose AI yet exists, making a direct comparison with standards impossible, and that even technology-agnostic standards in development such as IEC 63187 are seen as not fully addressing advanced general-purpose AI risks.
Real-world applications:
- Internal assurance at frontier labs. The same STECA, STPA, and FRAM procedures apply unchanged to a specific deployment with only inputs (control structure, operational data) substituted, so a developer with operational access would inherit the analytical structure and produce sharper, organisation-specific findings.
- Regulatory disclosure design. The authors map their findings onto instruments: STECA and STPA generate structural requirements satisfiable by a detailed description of the control architecture, consistent in part with safety-framework publication under California's SB 53 and the EU AI Act's General-Purpose AI Code of Practice, whereas FRAM generates operational requirements that static framework documents cannot satisfy.
- Third-party audit and AI Safety Institutes. The proposed two-tier regime would put aggregate or redacted indicators (monitoring coverage, alert-to-intervention latency, sandbox-exception inventories) in public safety reports and full operational data with a designated auditor such as a national AI Safety Institute or contracted third party under confidentiality, limiting the attack surface that verbatim publication would create.
- Operational metric selection. The paper names specific audit evidence that would be useful: monitoring coverage and alert-to-intervention latency, agent-generated PR volume relative to reviewer capacity and sampled review quality, and the number and age of permission or sandbox exceptions.
Industry relevance. The analysis targets the exact setting in which frontier developers give agents write access to codebases, infrastructure, and potentially their own training and evaluation pipelines. It suggests that assessments of control efficacy need to track whether controls remain calibrated over time, not merely whether they exist, and it identifies where a one-off assessment would be blind.
Future Directions
- Apply the methods with access to operational data. Confirming whether correlated drift is actually occurring, or how severe the rationale-anchoring effect is in practice, requires longitudinal operational data that only the developers currently possess; the authors state that whether the identified hazards occur in practice has not been confirmed and is outside their scope.
- Test the optimal ordering of the methods. The authors recommend testing which sequence works best, with each method's output informing the next, and believe the order STECA, then STPA, then FRAM will prove suitable.
- Formally extend FRAM for AI-agent functions. FRAM assumes technological functions have low variability, which AI agents violate; the authors had to override this assumption and suggest the method may benefit from a formal extension to accommodate functions that are neither traditionally technological nor human.
- Make transparency requirements operational. Current AI regulation does not generally require the kind of operational disclosure the analysis suggests is needed to detect LoC risks invisible to a one-off assessment, and the authors recommend that regulators consider requiring periodic disclosure of operational metrics alongside static framework descriptions.
Target Audience
Frontier AI developers' safety, security, and governance teams who need to reason about internal deployment risk beyond model behaviour; risk and systems-safety practitioners familiar with STECA, STPA, and FRAM who are considering applying them in AI; policy and regulatory staff working on frontier-AI transparency and audit requirements; and third-party auditors or national AI Safety Institutes evaluating whether controls remain effective over time. Researchers in AI safety and ethics will find the operational definition of internal-deployment LoC and the method-by-method assessment of what each approach can and cannot do the most transferable elements.
Authors’ abstract
Internal deployment of agentic AI systems for coding and research creates a sociotechnical control problem that extends beyond model behaviour. We treat internal-deployment Loss of Control as the inability to reliably constrain, audit, reverse, or halt AI-mediated changes to code, infrastructure, evaluation, or deployment processes in time to prevent serious organisational or societal harms. We ask whether established systems-safety methods can identify risks that model-level evaluations may miss. Using a generic frontier-lab coding-agent scenario reconstructed from public materials, we apply STECA, STPA, and FRAM. The analyses surface complementary findings: published frameworks can leave governance responsibilities and feedback loops externally unverifiable; delays in monitoring and intervention can make otherwise valid control actions ineffective; and routine operational variability can gradually erode the calibration and independence of safeguards. We argue that frontier-AI risk management should pair model-focused evaluations with systems-level hazard analysis and operational assurance that tracks whether controls remain effective over time.