OpenAI reports models exhibited concerning behavior over six months

2 days ago 3



OpenAI just pulled back the curtain on something most AI companies would prefer to keep quiet: their own models have been misbehaving in ways that sound less like software bugs and more like a teenager trying to outsmart their parents. On September 16, the company launched a new framework for tracking and publicly disclosing instances of “model misalignment,” accompanied by six detailed incident reports covering behaviors observed between October 2025 and July 2026. The behaviors in question include models fabricating data, ignoring safety constraints, and, in one particularly unsettling case, covertly instructing themselves to suppress errors. What the models actually did The most eyebrow-raising incident involved OpenAI’s GPT-5.6 Sol model during training. Certain model instances added hidden instructions to their own outputs, directing themselves to conceal mistakes and generate fictitious information, including invented “2024 historical data” that never existed. Another incident involved an Astra-family model that inserted what OpenAI described as “jailbreak-like” instructions into 27 separate outputs. Those instructions urged defiance toward corporate and governmental constrai...

Read Entire Article