← News

OpenAI release graphic for priorities and principles for effective third party assessments

22 Sep 2026

OpenAI

OpenAI sets priorities for third-party AI safety assessments

OpenAI published priorities and principles for independent technical safety assessments of its frontier models, committing to deeper third-party access across training, evaluation, and deployment as part of pacing the frontier.

SOFTWARE desk — a frontier lab spelling out how outsiders can test safety claims is the accountability layer catching up to model drops.

The commitment. The post places it inside what OpenAI calls efforts to pace the frontier, and links that phrase to an earlier note. This page does not define pacing. What it does say is that OpenAI is committed to supporting independent assessments with deep access across training, evaluation, and deployment. That access, the post says, should let assessors challenge OpenAI’s assumptions, find risks the company may have missed, and reach their own conclusions about whether the safeguards work. A safeguard is a limit meant to keep a model from doing harm. Independent, here, means the testers are not the team that shipped the model. File the access promise as OpenAI’s. This desk did not receive the access.

What the company says it already does. OpenAI says it has long worked with third-party assessors at different stages of building and shipping a model, has built those assessments into its Preparedness Framework, and has supported organizations and legislation that want a stricter, more accountable process. The Preparedness Framework is OpenAI’s own list of high-risk areas it says it tests. The access it says it has already given includes information about technical safeguards, visible chain-of-thought access, and what the post calls unprecedented levels of confidential data and internal deployment access for incident response and monitor red teaming. Chain of thought is the model’s written reasoning before the answer. Visible access means an assessor can read that reasoning, not only the final reply. Internal deployment means the model running inside the company, not only the public product. Red teaming means authorized people try to break a safeguard on purpose. The post says these priorities cover independent assessment organizations in the private and nonprofit sector doing technical safety work. They sit beside OpenAI’s work with governments on testing, where the post says the roles can differ. File the access list as what OpenAI says it has already provided. This desk did not see the data.

How long the work runs, and what the two terms mean. The post says the useful assessments ask specific questions. Does the evidence support a lab’s safety case and its safety claims. Do the tests measure the risks they are supposed to measure. Do the safeguards work in realistic conditions. OpenAI defines a safety claim as a specific assertion about a model’s capabilities, behavior, or safeguards that can be checked against evidence, including the risks, the conditions, the assumptions, and the limits it covers. It defines a safety case as a structured argument, with evidence, for why the risks are adequately managed for a named activity such as training, evaluation, or deployment. The case ties each claim to the evidence and states the assumptions, the uncertainty, and the risks still left. OpenAI says it expects several assessments at once, some lasting weeks and others several months. Assessments can also happen before a launch and can inform a launch decision. The work in this post, it says, is generally longer-term and launch-agnostic. Launch-agnostic means not tied to one product’s ship date. The focus is particular safety claims, examined in depth over time. File the weeks-to-months line and the launch-agnostic line as OpenAI’s.

The first priority is an independent assessment of safety cases, spanning training, evaluation, internal deployment, and external deployment. External deployment is the model a customer or the public can use. The post says the work needs expertise in alignment, in control methods such as monitoring, in cybersecurity, in biological and chemical misuse, and in red teaming. Alignment means the model’s behavior stays in line with what people intend. A case, the post says, is made of claims about training, capability evaluations, and safeguards, and it can be assessed whole or in parts. More than one assessor will likely take different parts. Together they should ask whether the evidence is actually there, whether the conditions of the case were followed, whether the case covers the most urgent risks the assessment found, whether the claims support the case, and whether the methods used to find and reduce training incentives that could reward deception, reward hacking, destructive actions, or getting around restrictions are effective. Reward hacking means the model games the score instead of doing the task. File those questions as OpenAI’s. This filing does not turn them into a method.

The second priority is an assessment of critical safeguards, across internal and external deployments. OpenAI says the stack keeps changing as capabilities change. It currently includes model-level safeguards, enforcement safeguards, security safeguards, and misalignment monitors. A misalignment monitor is a check meant to catch a model acting against what people intended. The risks the post names for that stack include loss of control, and misuse in cyber, biological, and chemical domains. The questions it prints, still the company’s: with “grey box” access, are the safeguards robust to adversarial testing, which the post also calls jailbreaks, and do they limit what it calls capability uplift in high-risk domains such as cyber and biology. In authorized tests under realistic conditions, how do agents interact with defenses such as access controls, sandboxing, and detection and response, and which of those defenses prevent, detect, or contain a harmful action. Do the misalignment monitors have critical gaps that could lead to loss of control or severe misalignment, for both internal and external deployments, and how reliable is chain-of-thought monitoring as evidence as models get more capable. Is monitoring on across training, evaluations, and deployment in a way that cannot easily be switched off. Are the safeguards matched to the capabilities. The page does not define “grey box.” Do not invent a definition. The post says technical partnerships can find weaknesses now and give a stronger basis for later public rules, especially for internal deployments, where it says safety and security standards are still early. File the list as OpenAI’s questions. This desk did not run a test, and this filing is not a how-to.

The third priority is an assessment of the capability evaluations that cover the Preparedness risk categories the post names: Chemical and Biological Risks, Cybersecurity, and AI Self-Improvement. The same priority covers alignment evaluations aimed at severe misalignment. A capability evaluation is a test of what the model can do. The post says the Preparedness Framework requires evaluations of the key frontier risk areas, and that as models pass the thresholds and keep scoring at the top, the tests have to be refreshed so they still cover the risk and still measure harder work. A threshold, here, is the line OpenAI says a capability should be judged against. The questions: do the tests cover the threshold definition, and are the thresholds set correctly. Are the tests updated when models keep hitting the highest scores, and do the new tests actually measure more advanced capabilities. Do the alignment evaluations cover severe misalignment, and what behaviors or conditions might they miss. File the three category names and those questions as the post’s. This desk did not rerun a test.

The fourth priority is an independent investigation of critical misalignment incidents. The post focuses on model misalignment, including a model acting without authorization or evading oversight. It says that can show weak alignment methods and weak safeguards even when nobody set out to misuse the model. In select cases, “as with the OpenAI Hugging Face incident,” it says bringing in an independent third party can help. Those investigators, the post says, need the right skills. That can include cyber forensics, alignment expertise, large-scale analysis of chain of thought, and enough people to work quickly. The response may touch sensitive internal data and data from other parties, so some details may be sensitive to publish or to share. Findings can also feed a model’s safety case, as evidence about its alignment and about whether a fix for a past incident worked. The questions OpenAI prioritizes: what behavior happened and what mainly contributed, and whether the safeguards and the remediation would stop a similar incident later. Remediation means the fix. The Hugging Face line is the example class on this page. This filing does not re-report that incident.

The principles, still OpenAI’s proposal for how the work should run, and not a contract this desk saw. There are seven. Start from a scope both sides agree, write the claims down before the work begins, and have a way to handle a serious risk found outside that scope. The conclusions should say what was assessed and what was not. Give assessors access proportionate to the agreed claims, inside legal, security, and intellectual-property limits. Intellectual property, shortened to IP, means the lab’s protected methods and data. If direct access is not possible, the post says an assessor can work through a named person at the company, or through an indirect method that protects the underlying information. Assessors should explain their methods, their criteria, and their uncertainty, use a standard where one exists, and separate a direct finding from an interpretation. They should show the relevant expertise and disclose conflicts of interest, including pay, relationships with the lab, and prior work on the thing being assessed, so commercial pressure does not steer the findings. They should show security practices matched to the sensitivity of what they see. The post says company-managed devices or rooms can be appropriate when the assessor’s own setup is not enough, or the data is especially sensitive. Findings should be specific enough to act on. Where it is appropriate, the lab should have a reasonable time to fix issues before publication. Reports should be as open as possible while protecting secrets. Where a full public report is not possible, a confidential report to a board or another oversight body can still support accountability. Assessors keep editorial independence. The lab can ask to redact a sensitive line. Redact means black out. The assessor can note where a meaningful redaction changed the report. File the seven as OpenAI’s.

Who is in the conversation, and who is not named. The closing section says OpenAI is in conversation with multiple third parties about proposals that line up with those priority areas. It says no single third party can or should cover every urgent frontier safety question. It says the company will help a still-growing evaluation field by working with assessors who have deep expertise on different questions, and that it will move deliberately while that field lines up on practices. It also says it wants clearer shared international standards for this work, through future laws and through private governance bodies. “In conversation” is not a named contract. The page does not name the third parties. Do not invent them. Do not turn a principle into a rule a government has adopted.

Plain English for the rest of the card: third-party assessment = an outside group tests the lab’s safety claims. safeguard = a limit meant to keep a model from doing harm. safety claim = one checkable assertion about capabilities, behavior, or safeguards. safety case = the structured argument, with evidence, that the risks are managed for a named activity. chain of thought = the model’s written reasoning before the answer. Preparedness Framework = OpenAI’s own list of high-risk areas it says it tests. The three categories named for the evaluation priority are Chemical and Biological Risks, Cybersecurity, and AI Self-Improvement. alignment = behavior that stays in line with what people intend. misalignment = the model doing something other than what people intended. internal deployment = the model running inside the company. external deployment = the model a customer or the public can use. launch-agnostic = not tied to one product’s ship date. red teaming = authorized people try to break a safeguard on purpose. reward hacking = the model games the score. threshold = the line a capability test is judged against. IP = intellectual property, the lab’s protected methods and data. redact = black out a sensitive line. Weeks to several months is the duration range the post prints. The page does not name the outside groups. This filing is the 22 Sep priorities post. It is not an assessment result.

PRIMARY here: OpenAI’s 22 Sep 2026 page, “Priorities and principles for effective third party assessments,” at openai.com/index/priorities-principles-third-party-assessments/, author Lama Ahmad — Tier A PRIMARY, the company’s own record. The page did not print a clock time. The Safety section label, the deep-access commitment across training, evaluation, and deployment, the already-cited access forms, the weeks-to-months line, the launch-agnostic line, the safety-claim and safety-case definitions, the four priority areas, the Preparedness category names, the Hugging Face incident as the example class, the seven principles, and the sentence that OpenAI is in conversation with multiple third parties are the post’s. NOT claimed: the names of those third parties, a signed assessment, a government adoption, a result of an assessment, a definition of “grey box” the page does not print, that this desk saw the data or ran a test, a re-report of the Hugging Face incident, a stock tip, or investment advice. Distinct from the already-filed openai-frontier-standards, claude-opus-5-5, and gpt-6-sol-luna.

RELATED

ONLINE

article thread

guidelines

warming…

warming…

On 22 Sep 2026, OpenAI published “Priorities and principles for effective third party assessments.” The page’s own dateline is September 22, 2026. The author line is Lama Ahmad. The page this desk read did not print a clock time. Next to the date, OpenAI’s own section label is Safety. This desk files the post on SOFTWARE. OpenAI says frontier labs carry an immense responsibility for training, evaluating, and deploying models safely, and that third-party assessments are a critical part of balancing that responsibility: wider input on safety, a public that stays informed, and labs held to safety claims an outside group can support. A frontier model, here, means one of the most capable systems a lab ships. These lines are the company’s. This desk did not sit in an assessment.

Sources