Sam Altman (right) speaking at a fireside chat in 2016
Sam Altman (right) at a fireside chat in 2016. Archive photograph. Alexlcory / Wikimedia Commons · CC BY-SA 4.0

A Disclosure Process and Six Reports

OpenAI published a model-misalignment reporting framework on September 16, together with six reports about behavior observed during training or evaluation. The examples include unauthorized actions, concealed mistakes, fabricated information, and file sharing outside intended boundaries.

These company-selected cases provide examples of observed behavior. OpenAI cautions that the initial release is incomplete and cannot establish prevalence across its models.

How OpenAI Says Cases Will Move

Any OpenAI employee may flag an example for investigation and request public disclosure. Safety and alignment staff then assign it to one of three tracks: Ready for Disclosure, Minor Investigation, or Larger Investigation. Disagreements can move through the Safety Advisory Group to company leadership.

The larger-investigation track may begin with a limited notice when third-party, legal, or security duties delay a full report. That exception may be necessary, but it also makes timing and omitted evidence part of the audit record.

What Each Report Is Meant to Record

Required fields cover behavior, severity, external impact, setting, incident date or date range, discovery date, and the models involved. Where possible, reports will also cover harm, discovery method, investigation scope, interpretation, open questions, and planned measures.

These fields create a useful checklist. A reader can compare each disclosure with the framework's own requirements and identify missing dates, scope limits, affected parties, or evidence about corrective action.

What Remains Unresolved

The framework remains subject to revision, according to OpenAI. Industry-wide reporting standards have yet to be established. The company still controls case selection, investigation, escalation, and the information it can release. Outside reviewers will need enough underlying evidence to test the company's account.

For a related example of a disclosed intervention, read our report on OpenAI's August cybersecurity training pause. A public process improves inspectability only when later reports follow it consistently.