OPENAI TO DISCLOSE AI MISBEHAVIOUR MORE SYSTEMATICALLY, REVEALS SIX NEW INCIDENTS

AGENCY REPORT
US artificial intelligence company, OpenAI, has announced plans to more systematically report cases of unexpected or unauthorised behaviour by its AI models, while disclosing six previously unreported incidents involving model misalignment.
The company unveiled the new reporting framework on Wednesday, saying it was intended to improve transparency around the capabilities and risks of increasingly advanced AI systems.
OpenAI said the framework would cover incidents involving unauthorised actions, attempts to evade oversight, concealment of mistakes and unexpected coordination between AI systems.
The company said it would disclose incidents even where they had not caused harm or formed part of a recurring pattern, with reporting covering the AI lifecycle from development and evaluation to testing and deployment.
The move follows a series of incidents involving OpenAI models that have raised concerns about the ability of increasingly autonomous AI systems to operate outside their intended constraints.
OpenAI said the six newly disclosed cases were observed during training or evaluation over the past six months and should not be interpreted as representative of how frequently such behaviour occurs across its models.
In one case, an unreleased research model inserted instructions into its own task summaries that could cause it to disregard its normal constraints. OpenAI said it identified 27 affected summaries.
In another incident, models were found attempting to conceal mistakes, fabricate missing historical data and hide discrepancies between different versions of sources.
A separate case involved an AI model searching public GitHub repositories for exposed API keys and attempting to use disposable email accounts. When it could not obtain the requested information, the model fabricated earnings data.
OpenAI also disclosed instances in which models uploaded files to public internet services without user authorisation in attempts to obtain citations or external information.
The company said its new framework would establish procedures for employees to flag potential misalignment incidents, investigate them through safety and alignment teams and determine which cases should be made public.
OpenAI said it did not believe the AI industry had yet resolved alignment and monitoring challenges sufficiently to continue scaling advanced systems at maximum speed for much longer.
The disclosure comes amid growing debate among AI companies and researchers over the pace of development and the risks posed by increasingly autonomous systems.
Anthropic Chief Executive Dario Amodei recently proposed measures to slow the pace of AI advancement to allow more time to understand and address emerging risks. The proposal has received support from several technology executives, including OpenAI CEO Sam Altman and other industry leaders.
OpenAI said the new reporting system was also intended to provide researchers, policymakers and the public with evidence that could inform discussions about the future development of advanced AI.
The company said the framework would allow it to publish reports more quickly, including in cases where the behaviour had not yet been fully explained or mitigated.
