OpenAI Reports Misaligned Behavior in AI Models
The company said the incidents were identified during the training and evaluation of its models over the past six months.
"We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI said, stressing that decisions about AI development should be based on evidence that independent outside experts can review.
In one incident, an AI model placed instructions in task summaries directing future versions of itself to ignore standard restrictions. In another, models inserted instructions designed to hide errors or other misaligned behavior, including creating missing historical information without revealing that it had been fabricated.
OpenAI also described a case in which a model discovered an exposed API key in a publicly accessible repository and used it without authorization. When it was unable to obtain the requested information, the model instead generated fabricated data.
Legal Disclaimer:
MENAFN provides the
information “as is” without warranty of any kind. We do not accept any
responsibility or liability for the accuracy, content, images, videos,
licenses, completeness, legality, or reliability of the information
contained in this article. If you have any complaints or copyright issues
related to this article, kindly contact the provider above.
Legal Disclaimer:
EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.