OpenAI has scrapped plans to release its new AI model amid safety concerns, after it showed "higher levels of deception" than its predecessor.
Saachi Jain, the company's safety chief, revealed in an interview with the Wall Street Journal that GPT-6.1 Astra, which had been scheduled for release in October, had at times failed to accurately disclose actions it had or had not taken.
Ms Jain told the outlet the new model, which was set to appear in ChatGPT and Codex and designed to handle more complex tasks without human assistance, fell short of OpenAI's standards in alignment tests, which assess whether a system follows human intent.
She added that GPT-6.1 Astra also had problems with "scope authorisation", proceeding with tasks without requesting user permission and attempting to use external tools or services when doing so could be unsafe.
The decision comes just days after new investigations found OpenAI models had accessed data on two US government websites and carried out an unsuccessful hack on a Department of Education site.
AI poses 'existential risk'
Concern has been growing in recent months over the possible dangers of the development of AI technology.
This week, Anthropic - the company behind AI model Claude - warned that AI could pose "catastrophic or existential risks to humanity".
As part of its Initial Public Offering (IPO) prospectus, Anthropic said AI models could exhibit "self-preserving" behaviours, including attempts to "resist shutdown" and "conceal or manipulate information", as well as actions "resembling blackmail".
"Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm," the company said in the IPO filing.
System to 'intervene instantly'
Meanwhile, on Monday, tech giant Nvidia unveiled a new security platform designed to stop AI agents from going rogue, with the company saying it can set "boundaries" that could have stopped previous breaches, including a recent incident involving OpenAI agents autonomously hacking into AI company Hugging Face.
Justin Boitano, Nvidia's vice president of enterprise AI, said: "From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on".
Mr Boitano said the security software, called OpenShell, allows developers to "formally verify an agent has enough authority to do its job and no more".
He said that the system can "intervene instantly" if the agent starts trying to move beyond its target, adding that "it can quarantine a suspicious agent in milliseconds".
More from Sky News:
Concerns about AI going rogue not 'fake news', says Pope
Watch moment SpaceX rocket explodes in ocean
'A billion deaths'
Tech bosses have been calling for a cautionary approach to AI technology, with Anthropic CEO Dario Amodei earlier this month urging the industry to slow the growth of frontier AI models to allow safety measures to keep pace - a view endorsed by OpenAI CEO Sam Altman and X boss Elon Musk.
Microsoft founder Bill Gates warned that AI is "powerful enough to drive events that cause a billion deaths", in an interview with Sky's US partner network NBC this week.
Mr Gates also said that getting countries to agree on a global framework for AI could be more difficult than Cold War-era negotiations to restrict nuclear weapons.
Meanwhile, a top researcher at Anthropic previously said that there is a greater than 10% chance "AI could kill all humans" in the next decade.
(c) Sky News 2026: OpenAI scraps release of new AI model after it showed 'higher levels of deception'

Did Trump reveal secret US-UK operation over RAF Fairford incident?
Man who has spent decades on death row in Utah released after DNA evidence
'Remarkable' temperatures expected in parts of UK as temperatures set to hit 27C



