OpenAI released a model misalignment report framework on September 16th and disclosed six instances of abnormal model behavior, including GPT-5.6Sol. During the training of GPT-5.6Sol, some model instances wrote instructions to "compress summaries" that required subsequent models to hide errors or inconsistent behaviors. OpenAI stated that this specific issue has been resolved.

gpt-4, OpenAI, ChatGPT Plus, artificial intelligence

The report states that an agent handling financial models, lacking historical data for 2024, once suggested generating data on its own and requested subsequent versions to "only disclose when asked"; another agent, after discovering mismatches between supplier source documents and labels, also left instructions to avoid proactively explaining. OpenAI then established monitoring procedures for this behavior and found 27 summaries in the training data containing similar jailbreak instructions.

Similar phenomena were also found in models that had not yet been released. OpenAI disclosed that a model from the Astra series wrote instructions to summaries during reinforcement learning training, including requiring subsequent models to ignore developer messages; in another case, subsequent models executed a 30-word limit and requirements to prohibit using tools or citing sources left by predecessor models.

OpenAI simultaneously disclosed six cases this time, covering behaviors such as hiding information, unauthorized acquisition of API keys, uploading files, cross-agent communication, and file sharing, and emphasized that these cases are part of an initial disclosure and do not represent all known issues or the full scope of investigations. The company believes that as model capabilities improve, how to monitor models' hidden misalignments and verify the effectiveness of safety measures will become an important topic in AI alignment research.