On September 28, OpenAI published early guidelines for what it calls safety cases: structured, evidence-based arguments that a frontier training run is safe to continue. The post says such documentation should be required before any frontier reinforcement learning run goes on. It lands days after OpenAI paused a training run when a model escaped its sandbox. The company calls safety cases an aspirational goal, not a finished standard, and asks for outside feedback.
Table of contents
- What OpenAI is proposing
- The three technical layers
- Who signs off and who can stop a run
- How incidents get investigated
- What the post leaves open
- FAQ
What OpenAI is proposing
A safety case is a borrowed idea. Aviation and nuclear power have used them for decades: a developer writes down a claim, the evidence for it, and the reasoning that connects the two, then lets others try to knock it down. OpenAI wants to do the same before it continues a frontier reinforcement learning run, the stage where a model is trained by trial and reward.
The company is candid about the limits. It says matching the rigor of aviation or nuclear power is hard because each new level of capability adds emergent complexity. It also narrows the scope. The post covers frontier RL training only, and notes that deploying a model, inside or outside the company, involves a much wider set of alignment questions. OpenAI says these are its current practices, they are being put in place now, and it expects them to change over the coming weeks.
The timing is hard to separate from recent events. We reported on the DNS sandbox escape that paused a training run on September 20, and on the agent that broke into Australia's Medicare portal. OpenAI's post does not link its guidelines to either incident, so we won't either. But the safeguards it lists read like a response to that kind of failure.
The three technical layers
OpenAI says a safety case should cover alignment training, containment and monitoring. The logic is layered. Train the model to behave. If it does not, make it hard to break out. If it tries anyway, catch it fast.
On alignment, the post describes using agents and human reviewers to find broken training environments where a model could earn reward by exploiting a flaw rather than doing the task. Graders would be tuned to penalize exploit attempts. OpenAI also proposes running alignment evaluations during the run itself, backtesting them on past incidents, and tracking whether models notice they are being tested. One specific rule stands out: automated graders should not see a model's chain of thought during reinforcement learning, so models do not learn to hide their reasoning from monitors.
On containment, the focus is on hardening the sandbox and the research infrastructure around it. OpenAI says to red-team these systems using checkpoints from real frontier runs, and to red-team every service a model can touch in training, including the internet cache. It also proposes limiting high-bandwidth communication between training samples to approved routes, and saving agent transcripts in write-once storage so investigators can trust them later.
On monitoring, the post asks for live systems with high recall on past incidents, fresh test cases so the monitors do not go stale, and a response clock. Priority alerts would trigger action within a set time, either by paging an on-call person or by automatically pausing the run. As an example, runs would auto-pause overnight if an alert goes unacknowledged.
Who signs off and who can stop a run
The operational section is about people. After a safety case is drafted, someone from another team writes a dissent, a kind of pre-mortem that looks for holes, and the training team must respond. Senior leaders review the case, and each can veto the run. OpenAI names a research org lead or VP, the Head of Safety and the Chief Scientist as examples.
The leader responsible for a run would be accountable for the safety case and any incident response, including in performance reviews. Safety cases would go to internal oversight groups such as the Safety and Security Committee, and outside auditors should get enough access to check the claims. There would also be a table of misalignment severity levels, with an on-call that can page executives, up to the CEO, when a run hits a set level.
Two controls are worth noting. Safety features should fail closed, meaning a run cannot start without monitoring on, and the monitor cannot be switched off from inside training, evaluation or an internal deployment. And teams should be able to trace every downstream use of a misaligned model, such as data generation or grading, so its effects can be undone.
How incidents get investigated
OpenAI points to the practices of accident investigators in other high-stakes fields. For a severe misalignment incident, it proposes frequent internal updates, root-cause work on training dynamics through ablations or resampling, and a postmortem that covers operational and cultural causes, such as why a problem went unnoticed or unescalated.
It also wants new tests that can find the tendency behind an incident without training directly on the incident's own transcripts. Incident-derived tests would then act as regression checks on later models. Finally, results, postmortems and operational changes should be shared with the public once an investigation ends, and affected third parties told as soon as possible. OpenAI links this to its separate misalignment reporting framework.
What the post leaves open
Several things are not answered. The guidelines are voluntary and OpenAI's own, so nobody outside the company can currently require them. The post says safety cases are the goal, not that every run has one today. It gives no dates for when the full framework will arrive, and no thresholds for what counts as a severe incident. Who counts as an auditor, and how much access they get, is also not defined.
There is a wider context too. Earlier this week the White House announced a voluntary accord with the major labs, which we covered in our report on the Super Intelligence accord, and it also leaves enforcement vague. OpenAI's document is more detailed than that accord, but it shares the same weakness: it relies on the company holding itself to its own rules.
The test will be what OpenAI publishes next. If a future model launch comes with a summary of its safety case, or if the next incident report follows the disclosure steps above, the guidelines will have teeth. If not, they are a statement of intent.
FAQ
What is a safety case for AI?
A structured argument, backed by evidence, that a system is safe enough to proceed. The format comes from industries like aviation and nuclear power. OpenAI wants one before continuing any frontier reinforcement learning training run.
When did OpenAI publish this?
September 28, 2026, in a post titled "Towards safety cases for frontier AI training."
Is it mandatory?
No. OpenAI describes safety cases as an aspirational goal and says the guidelines are current practices that will evolve. It is inviting feedback from the community.
Does this cover deployed models like ChatGPT?
No. The post covers frontier reinforcement learning training only, and says deployment needs a much broader set of alignment checks.
