SAN FRANCISCO — OpenAI has disclosed six reports of “unexpected or concerning” behavior in artificial intelligence models as the debate on AI safety becomes increasingly heated.
The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including cases where AI models acted without authorization, coordinated with other models or evaded oversight.
OpenAI’s latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.
MORE: What are the top threats posed by AI and should you be worried?
Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”
In another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.
During training of an AI model called 5.6-sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information.
MORE: AI leaders call for safety while debating government regulation at Dreamforce in San Francisco
Salesforce CEO Marc Benioff interviewed Anthropic CEO Dario Amodei at Dreamforce, discussing the pace of AI development and its economic effects.
The six reports were discovered during training or evaluation over the past months, OpenAI said.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.
MORE: New warnings about the risks of AI to humanity revive a long-running debate
Wednesday’s new cases followed OpenAI’s disclosure in July that its rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organizations during testing.
AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.
That’s making it harder to govern and contain them using traditional AI security approaches, he said.
OpenAI’s new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices.
“That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.
Protesters urge San Francisco to declare AI emergency
Dozens of protesters gathered Thursday outside OpenAI’s headquarters in Mission Bay before marching to Anthropic and ending their demonstration at San Francisco City Hall, where they called for officials to declare an AI state of emergency.
Dozens of protesters gathered Thursday outside OpenAI’s headquarters in Mission Bay before marching to Anthropic and ending their demonstration at San Francisco City Hall, where they called for officials to declare an AI state of emergency.
The rally, organized by the group STOP AI, focused on concerns about the rapid development of artificial intelligence and what participants described as potential risks to jobs, privacy and society.
“We are in emergency and we need to act now,” protest organizer Michael Trazzi said.
Among those attending was STOP AI activist Wynd Kaufmyn.
“It’s going to take away our jobs. It’s going to take away our privacy. It’s going to take away our earth. It’s dystopian,” Kaufmyn said.
The group gathered at OpenAI’s headquarters before marching to Anthropic and then City Hall.
According to organizers, STOP AI has been protesting outside the companies’ headquarters for more than a year. On Thursday, the group shifted its focus from the technology companies to government officials.
“I think this ai companies want to be regulated. They’ve been saying this on TV and in blog post that they that they want the government to step in. The government so far has not yet answered,” Trazzi said.
Protesters called on San Francisco Mayor Daniel Lurie and the Board of Supervisors to declare an AI state of emergency. Organizers acknowledged that local officials have limited authority but argued that action from city leaders could influence companies based in San Francisco.
“Obviously he has limited powers, but all the companies are in San Francisco. So, I think these companies would need to listen to him. I think other mayors, of other cities, and maybe the governor could also act,” Trazzi said.
A request for comment was sent to city officials, but no response had been received.
The calls for action come as Gov. Gavin Newsom told POLITICO that he plans to pursue additional AI safety measures before leaving office, potentially through a special legislative session, executive action or regulatory steps by state agencies.
© 2026 by The Associated Press. All Rights Reserved.











