← News

OpenAI publishes framework for reporting model misalignment

OpenAI said it is sharing a new framework for tracking, investigating, and disclosing instances of model misalignment, along with six reports on unexpected or concerning model behavior from the last six months. The company says prior disclosures were ad hoc and less frequent than ideal, and that the new process is meant to publish reports faster — even when behavior is not yet fully explained or mitigated. Cases go through tracks including Ready for Disclosure, Minor Investigation, and Larger Investigation (Slow Track); OpenAI says the Hugging Face incident would have fallen under the Slow Track had this framework existed then.

SAFETY desk — same-day primary that a frontier lab is turning ad hoc misalignment write-ups into a standing public disclosure process, and is pairing it with a fresh batch of incident reports — evidence outsiders can examine while the industry argues about pace and oversight.

What the company says it is sharing: a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. Misalignment here means a model doing something unexpected or concerning — behavior that is off the intended goal — not a finished proof that the model has “gone rogue.” File that tracking / investigating / disclosing picture as OpenAI’s. This desk is not inventing a law, a regulator mandate, or an industry-wide standard that other labs have already signed.

OpenAI also shared six reports on unexpected or concerning model behavior observed in the last six months. File the six-report count and the last-six-months window as OpenAI’s. This desk is not inventing the six incident titles, a severity ranking, or a claim that those six are the only cases in that window.

Why the company says it needed a process: prior misalignment disclosures were ad hoc and less frequent than ideal. OpenAI says it often waited to collate several instances or put them in system cards. A system card is a lab write-up that ships with a model and describes how it was tested and what risks showed up. File that ad hoc / less-frequent / collate-or-system-card picture as OpenAI’s. This desk did not audit older system cards as a census.

Goal, still company: expedite publishing misalignment reports after observation, even when OpenAI has not fully explained or mitigated the behavior. File that publish-faster / even-if-unexplained picture as OpenAI’s. This desk is not inventing a day-count SLA — a promised number of business days — unless the OpenAI primary page prints one. None is used as a lead fact here.

Pace line, still OpenAI’s: the company says it does not believe the AI industry has solved alignment and monitoring enough to continue responsibly scaling at maximum speed for much longer. Alignment here means steering a model so it does what people intended. Monitoring means watching what the model actually does. File that unsolved-alignment / not-maximum-speed line as OpenAI’s. This desk is not treating it as a pause, a halt, or a signed industry pact.

Disclosure tracks, as the same post has it: cases go to Ready for Disclosure, Minor Investigation, or Larger Investigation, which OpenAI also calls the Slow Track. Ready for Disclosure and Minor Investigation are expected to cover the large majority of disclosed instances. Today’s six, the company says, fall into those two tracks. File those three track names, the majority line, and the six-in-the-first-two-tracks line as OpenAI’s. This desk did not assign a case to a track.

Slow Track, still company: Larger Investigation covers complex cases, especially involving third parties. Security, legal, and responsible-disclosure obligations can delay notices. Responsible disclosure here means telling an affected party before, or instead of, dumping every detail in public. OpenAI says the Hugging Face incident would have fallen under the Slow Track had this framework existed then. File that third-party / delay / Hugging-Face-would-have-been-Slow-Track picture as OpenAI’s. This filing is the framework plus the six-report batch — not a re-investigation of Hugging Face.

Same-day independent wires confirm the publish: Wired and Axios both reported the framework and the six-report batch on 16 Sep 2026. File those as Tier B confirm of the same-day announce — not as a substitute primary, and not as a source for invented SLAs or invented incident titles. Extra wire color — including named examples Wired attached to some of the six — stays in Sources.

Plain English for the rest of the card: misalignment = a model behaving in an unexpected or concerning way, off the intended goal. system card = a lab write-up that ships with a model and describes tests and risks. Ready for Disclosure = the fast public-report track. Minor Investigation = a short extra look before publish. Larger Investigation / Slow Track = the slower track for complex cases, especially with third parties. responsible disclosure = telling an affected party before dumping every detail in public. Hugging Face incident = the already-filed July agent-misalignment case OpenAI says would have been Slow Track under this process. six reports = the company’s same-day batch from the last six months — titles not invented here.

PRIMARY here: OpenAI’s 16 Sep 2026 index post “Our framework for reporting model misalignment” — Tier A PRIMARY company source, the original record. Wired and Axios are same-day independent Tier B confirm of the framework plus six reports, not a substitute primary. The Hugging Face incident hub is prior-cycle context, not today’s originating announce. The new tracking / investigating / disclosing framework, the six-report count from the last six months, the ad hoc / less-frequent / collate-or-system-card past, the publish-faster-even-if-unexplained goal, the unsolved-alignment / not-maximum-speed line, the Ready for Disclosure / Minor Investigation / Larger Investigation (Slow Track) tracks, the majority-in-the-first-two-tracks line, today’s six falling in those two tracks, Slow Track delays for third-party / security / legal / responsible-disclosure cases, and the Hugging Face incident as a would-have-been Slow Track case are company-attributed. NOT claimed: a day-count SLA, the six incident titles, that other labs adopted the framework, a U.S. legal duty to publish, that this desk verified each report, a pause or halt, a stock tip, or investment advice. Distinct from the already-filed openai-rogue-agents-ten-more-sites, openai-agents-rubygems-gemstuffer, anthropic-alignment-assessment-cyber-incidents, china-tc260-ai-safety-framework-3-0, lawzero-300m-canada-germany, amodei-pace-the-frontier, and claude-cowork-chat-merge.

RELATED

ONLINE

article thread

guidelines

warming…

warming…

On 16 Sep 2026 OpenAI published “Our framework for reporting model misalignment.” The company PRIMARY is the OpenAI index post at https://openai.com/index/model-misalignment-reporting-framework, dated September 16, 2026. That company research/safety post is the filing event. These are OpenAI’s words. This desk did not sit inside OpenAI’s safety process, and it did not independently verify each of the six reports.

Sources