OpenAI discloses 6 new cases of ‘misaligned’ AI behavior

3 hours ago 2



OpenAI on Wednesday disclosed another six cases of “unexpected or concerning” model behavior over the last six months.In a blog post, OpenAI said the cases illustrate a range of different behaviors it classifies as “misaligned behavior,” such as concealing information from the user and taking “unsanctioned actions” to overcome obstacles. The disclosures add to concerns among AI developers and researchers about whether safeguards are keeping pace with increasingly capable models. Last week, Anthropic CEO Dario Amodei called for a slowdown in frontier AI development, warning that unchecked AI advancement may “outrun our ability to understand and control these systems.” OpenAI said its disclosures were made to “inaugurate” its new framework for reporting model misalignment, and the cases shouldn’t be considered reflective of how often misalignment occurs across its models. According to OpenAI, one instance saw an “unreleased research model” insert “jailbreak-like instructions” in its own task summaries (used when continuing a task in a new context window), such as ignoring developer messages or adopting an unrestricted persona. Researchers found 27 summaries containing such instructio...

Read Entire Article