OpenAI identifies six additional safety concerns and unveils incident-disclosure plan

OpenAI has disclosed six more safety issues affecting its systems and introduced a new framework to log, investigate and publicly report instances of model misbehavior, or 'misalignment'. The move aims to increase transparency around how the company handles safety failures.

OpenAI said it had identified six additional safety issues affecting its models and outlined a new process to monitor and disclose cases in which its systems behave in unintended or unsafe ways.

The company described the initiative as a system to track, investigate and publicly report incidents of what it terms "misalignment" — situations where model outputs diverge from intended behaviour or safety expectations. OpenAI framed the step as part of broader efforts to make its safety work more systematic and visible to external audiences.

Officials at the firm said the disclosure mechanism is intended to create a more accountable record of safety failures and of the steps taken to address them. The plan covers the lifecycle from detection and investigation through to disclosure, enabling outside observers to see what issues occurred and how the company responded.

Industry observers have increasingly called for clearer reporting on AI safety incidents as models are deployed more widely. OpenAI's announcement follows that trend, signalling an attempt to provide greater transparency about the risks its systems can pose and the mitigation measures it applies. The company did not provide further technical detail about the six newly reported issues in its initial statement.