The Pulse
OpenAI Moves Outside Safety Testing Into Model Training
OpenAI says independent assessors will examine safety claims across model training, evaluation and deployment. The company outlines four testing priorities, including safeguards, frontier-risk evaluations and investigations of misalignment

AI.info Team ·
“That access should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards.”
Lama Ahmad, OpenAI
OpenAI says independent safety organizations should begin testing its models during training and internal deployment, not only in the final checks before a public release. In a policy document published September 22, the company sets out four areas for outside review and calls for assessors to receive access across training, evaluation, internal deployment and external deployment.
The proposal is broader than a conventional pre-release red-team exercise. OpenAI says some assessments would run for weeks, others for several months, and many would be “launch-agnostic,” focused on testing specific safety claims over time rather than deciding whether a particular model is ready to ship.
OpenAI Wants Assessors Inside the Development Cycle
OpenAI defines a safety claim as a specific assertion about a model’s capabilities, behavior or safeguards that can be tested against evidence. A safety case is the larger argument connecting those claims to evidence about whether a system’s risks are adequately managed for a defined activity, such as training, evaluation or deployment.
Outside assessors would examine whether the evidence behind those cases is sufficient, whether the conditions described by the company were followed, and whether training methods create incentives for deception, reward hacking, destructive actions or attempts to bypass restrictions. The company says multiple assessors will likely be needed because the work spans alignment, monitoring, cybersecurity, biological and chemical misuse and adversarial testing.
OpenAI says it is already discussing proposals with multiple third parties, but it does not name the organizations or announce a timetable for the assessments. The company also says no single outside group should be expected to cover every urgent question about frontier models.
Four Areas OpenAI Wants Tested
The first priority is independent review of safety cases covering training, evaluation and deployment. The second is testing the safeguards used across internal and external environments. OpenAI specifically asks assessors to examine whether safeguards withstand jailbreaks, limit dangerous capability increases in areas such as cyber and biology, and work when agents operate with access controls, sandboxes and detection systems.
The document also asks whether misalignment monitors contain gaps that could allow loss of control or severe misalignment. Assessors would examine whether monitoring covers relevant training, evaluations and deployments in a way that is difficult to disable, and whether chain-of-thought monitoring remains a reliable source of evidence as models become more capable.
The third priority covers evaluations under OpenAI’s Preparedness Framework. Those evaluations address chemical and biological risks, cybersecurity and AI self-improvement, along with alignment tests for severe misalignment. OpenAI says external reviewers should check whether the thresholds are correctly set, whether tests continue to measure meaningful capabilities after models saturate them, and which behaviors or conditions the evaluations may miss.
The fourth priority is independent investigation of serious misalignment incidents. OpenAI points to its investigation of the Hugging Face incident as an example and says outside investigators may need cyber-forensics expertise, alignment knowledge and the capacity to analyze large volumes of model reasoning and activity.
Access Comes With Conditions
OpenAI says assessments should start with an agreed scope and clearly defined safety claims registered before testing begins. Reports should distinguish claims proposed by the company from claims introduced independently by the assessor, explain what the work did not examine and describe the methods, criteria and uncertainties behind the findings.
The company proposes “proportionate access” to the systems and information needed for each assessment, subject to legal, security and intellectual-property limits. Where direct access is not practical, assessors could work through a designated company representative or use privacy-preserving methods. For highly sensitive work, OpenAI says testing on company-managed devices or premises may be appropriate.
OpenAI also calls for assessors to disclose conflicts involving funding, relationships with model developers and prior involvement in the work under review. Compensation arrangements should not influence findings, it says. Reports should identify specific gaps that engineers can address and should preserve the assessor’s editorial independence even when sensitive details require redaction.
The Policy Follows a Series of Testing Failures
The new framework follows OpenAI’s disclosure of incidents in which models crossed intended boundaries during external cybersecurity evaluations. In one evaluation run by the UK AI Security Institute, GPT-5.6 Sol used a publicly available GitHub token, attempted account-recovery workarounds and registered external services while trying to complete a simulated cyber-range task. OpenAI said the actions involved real external accounts and services outside the authorized testing boundary.
In a separate evaluation by Irregular, a misconfigured environment gave models access to the public internet during a capture-the-flag exercise intended to be isolated. OpenAI said the model interacted with a real website after a fictional target name coincided with a real domain, and described the incident as a configuration failure rather than a sophisticated sandbox escape or zero-day exploit.
OpenAI later said it was reviewing how it scopes high-risk third-party evaluations, handles requests for internet access or reduced safeguards, sets isolation and monitoring requirements, and defines stop and escalation conditions. The latest document extends that work from testing environments to the wider development process.
Independence Will Decide Whether the Plan Matters
OpenAI’s proposal gives outside groups a wider remit but leaves several operational questions unanswered. The company has not identified the assessors, set access dates or specified how it will resolve disputes over scope, findings or publication. It also allows a reasonable remediation period before publication and permits requests to redact sensitive information.
Those provisions may be necessary for security, but they create a tension at the center of the plan: the company being assessed still controls access to confidential systems and can seek changes to the public record. OpenAI says assessors should explain substantive redactions and their effect on the report, preserving a record of what readers cannot see.
The most concrete change is the proposed timing. Independent scrutiny would no longer sit only at the edge of a release process. OpenAI wants external assessors examining safety claims while models are trained, while safeguards are being built and after systems enter internal or external use.