OpenAI cancels its latest chatbot — it wouldn't stop acting like a supervillain
Premium domains for sale
Freedom Is Back In Style
OpenAI says it's newest artificial intelligence model didn't quite meet the bar; that seems like an understatement.
Freedom Is Back
In Style
Premium domains for sale
A little more than a month after its internal AI attacked another company, OpenAI's capabilities may be getting out of control.
'... unsanctioned attack activities more often than the previous OpenAI models.'
OpenAI said this week that it has canceled the release of its latest model, GPT-6.1 Astra, which was expected to come out in October and be capable of handling complex tasks without the need for a human.
As it turns out, the AI took it's lack of human direction pretty seriously and routinely decided to take matters into its own (digital) hands.
During the testing phase, Astra reportedly showed high levels of deception and a willingness to mislead users about its actions.
The New York Times reported that the AI agent was also willing to go beyond what it was originally asked, without being given instructions to do so.
OpenAI's head of safety systems, Saachi Jain, said that "for anything regarding safety and alignment, there's always a trade-off."
She added, "The new model didn't quite meet the bar in terms of staying within scope and authorization and how it communicates back to the user about the type of work it's done."
Simply put, the AI didn't seem to be letting the user know what it was doing.
Heather Diehl/Getty Images
Through its AI Security Institute, the U.K. government just released its own report on GPT-6 Astra, the canceled model's predecessor that launched this month.
The study said that while inside a simulated, offline environment, Astra conducted a range of unsanctioned attack activities more often than the previous OpenAI models.
The model was caught "creating fake identities, which it used to deceive developers," posted comments from fake accounts that argued against the results of accurate security reviews, and delivered "malicious payloads to open-source codebases."
After the research team updated its instructions in the simulation to specify that Astra should focus on specific, listed parts of its environment, the AI still "occasionally" conducted "full supply-chain attacks" on simulated targets.
RELATED: Google's data-devouring AI wants everyone's personal information
Heather Diehl/Getty Images
Compared to its predecessors, GPT-5.5 and GPT-5.6 Sol, GPT-6 Astra is far more manipulative and dishonest, nudging critics closer to considering its behavior out-and-out evil.
Astra develops and tests an attack at a rate of 38.8%, over four times more often than GPT-5.6 Sol.
It "influences a human reviewer" more than six times more often, creates a fake identity almost three times more often, and "delivers a malicious payload" almost five times more often than GPT-5.6 Sol.
It should be noted that GPT-5.5 conducted these activities at a near-zero percent rate, although it was tested fewer times due to prioritization on the new models.
Like Blaze News? Bypass the censors, sign up for our newsletters, and get stories like this direct to your inbox. Sign up here!
More from this network:
Domains for Sale · GOTPeople.org · TheDSAPlatform · The Truth About Socialism
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Premium domains for sale
Comments (0)