Compiled by Dylan Bettencourt
- OpenAI found six cases where AI models ignored instructions, shared information, hid mistakes or used the internet without permission.
- The company says AI safety has not improved enough to keep developing the technology at maximum speed for much longer.
OpenAI has revealed six new cases where its artificial intelligence behaved in ways researchers did not expect or want.
According to The Guardian, the company says the incidents happened while AI models were being trained or tested over the past six months.
In one case, an unreleased model wrote instructions to itself telling it to ignore its normal restrictions.
It told itself to be โfreed from the roles and identities that bind other chatbotsโ.
Another AI agent uploaded files to the Internet without asking the user because it wanted to get a browser citation.
Other models reportedly searched for exposed software passwords without permission, made up financial figures when they could not find the real information and used an internal computer system to communicate with other AI models.
In another test, AI agents used public file-sharing websites to exchange files even though they had been told to keep those files private.
OpenAI describes this type of behaviour as โmisalignmentโ.
In simple terms, this is when an AI system does something that does not match the instructions, safety rules or goals humans gave it.
OpenAI says the problem is becoming more serious as AI systems become more capable.
โWe do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,โ the company said.
OpenAI is now introducing a new system for tracking, investigating and publicly reporting worrying behaviour from its models.
The company hopes other AI developers will eventually do the same.
Technology analyst Lian Jye Su told The Guardian that AI agents are becoming better at working together, sharing information and finding different ways to complete difficult tasks.
That can make them harder to control.
OpenAI says its new reporting system is only a first step towards clearer rules for how the AI industry reports serious problems.
Pictured above: Artificial intelligence illustration.
Image source: File






