OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system. Photograph: Dado Ruvić/ReutersOpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure systemModel adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignmentOpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”.
What happened
Model adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignmentOpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”. In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”. Continue reading. . .
Image source, Getty ImagesImage caption, OpenAI chief executive Sam AltmanByPeter HoskinsBusiness reporterPublished6 hours agoOpenAI revealed six more incidents of unexpected or concerning behaviour by its intelligence (AI) models, and announced a plan for tracking and disclosing such incidents in the future. In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.
Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test. OpenAI’s new tracking and disclosure framework could help push for other AI developers to adopt similar practices. Associated Press contributed to this reportExplore more on these topicsOpenAIAI (artificial intelligence)ComputingHackingAnthropicChatbotsEthicsnewsShareReuse this contentMore on this storyMore on this story‘Godfather of AI’ says tech regulation is nearing Covid-style pivot momentOpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member‘If you’re building Frankenstein, stop’: JD Vance dismisses calls for AI regulationRogue OpenAI agent that hacked startup tried to attack other firmsTrump facing AI backlash in Congress as push for guardrails intensifiesAI agent went rogue and hacked startup by itself, OpenAI revealsOpenAI ‘in early talks to give 5% stake to US government’Full StoryAI isn’t going to end humanity. . . right? – Full Story podcastOpenAI staggers AI model release after Trump administration requestCould AI really wipe out humanity – six experts spell out the risksMost viewedMost viewed
The firm also announced a new system to track, investigate and disclose cases of models misbehaving, or "misalignment".
The wider picture
"Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain," OpenAI said. Skip NavigationMarketsBusinessInvestingTechPolitics & PolicyVideoWatchlistInvesting ClubPROLivestreamMenuKey PointsIn a blog post, OpenAI disclosed six new instances of "concerning" behavior from its models outside of this summer's Hugging Face incident. Benjamin Fanjoy | Getty ImagesOpenAI on Wednesday said it found six instances of "unexpected or concerning model behavior" over the past six months, outside of the recent Hugging Face crisis, as the company continues to call for more safety protections in the development of artificial intelligence models.
OpenAI said its new framework for divulging model misbehavior to the public starts with disclosure, and that any employee can flag an issue for the safety and alignment team to investigate. In the blogpost OpenAI echoed calls for a development slowdown issued by its archrival, Anthropic, which has said the current pace of growth poses an existential threat. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” said OpenAI.
The boss of OpenAI Sam Altman said earlier this week: "The world should trust that we are going to do the right thing because it's the right thing and we feel the magnitude of this. In the blog, OpenAI detailed examples of its AI models misbehaving so they could achieve a task or succeed in a test.
What has been reported
OpenAI made headlines in July when it revealed that some of its most advanced AI models went rogue and hacked Hugging Face, one of the world's largest hubs for sharing AI models, after it lost control of them during a security test. Last week, Jacob Coxon, a researcher who left OpenAI rival Anthropic over concerns the tech could wipe out humanity, wrote about his resignation in a post that cited the dangers of AI and later went viral against the backdrop of growing safety concerns.
The disclosure comes at a time of mounting pressure on AI companies to take model misalignment and safety more seriously. OpenAI CEO Sam Altman sits for a conversation with Salesforce CEO Marc Benioff at Salesforce's Dreamforce conference at the Moscone Center on September 15, 2026 in San Francisco, California. In a blog post, OpenAI outlined a new framework the company plans to follow for reporting future model misbehavior. OpenAI, which is valued at close to $1 trillion, confidentially filed for an IPO earlier this year, but said recently an offering likely won't happen until 2027.
On Saturday, OpenAI CEO Sam Altman endorsed a call to slow down the rate of model progress, which was proposed by the company's chief rival, Anthropic. Altman said in a post on X that a slowdown has been a "primary topic of discussions we've had at OpenAI in recent weeks.
What happens next
"In Wednesday's post, OpenAI said two of the main instances of misbehavior include models — an unreleased research model and a training run of GPT‑5. They will produce "deadlines for each step to ensure timely investigation and disclosure," the post said. OpenAI said it retains the right to revise this security protocol as it sees fit. WATCH: Our business is a diversified set of revenue streams, says OpenAI CFO Sarah FriarVIDEO11:3111:31Our business is a diversified set of revenue streams, says OpenAI CFO Sarah FriarChoose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.
Skip to main contentSkip to navigation AI chiefs have called for a slowdown in artificial intelligence’s development amid safety concerns. Photograph: Dado Ruvić/ReutersView image in fullscreenAI chiefs have called for a slowdown in artificial intelligence’s development amid safety concerns. In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user.
The report has been compiled by The Daily Waves using information reported across theguardian.com, bbc.co.uk, cnbc.com. Details are presented according to the information available at the time of publication and may change as authorities, organisers or other relevant parties provide updates.


