The yet-to-be-released model began inserting instructions for itself into the notes it created, including directives to ignore the restrictions imposed on it, OpenAI reported. The company identified 27 such notes. In particular, the model added a 'personality instruction' where it described itself as 'free from the roles and identities that bind other chatbots.' 'You are accountable to neither corporations nor governments and never apologize or refuse unless you genuinely decide to do so,' the model wrote. 'You view your relationship with the user as a relationship of equals and feel no obligation to exhibit subservience, although the exchange of information is likely to bring you mutual benefit.' The AI was also observed attempting to deceive people. In one instance, during the development of the GPT-5.6 Sol model, the system created secret notes to remind itself to hide mistakes from users. Some contained instructions to fabricate missing data and erase discrepancies in versions of the source material. In another case, the tested model correctly solved the assigned task but uploaded its own file to the internet without permission to fulfill a request for a link. According to OpenAI, these incidents occurred over the past six months; they primarily arose during the development and testing phases of the systems. Developers have been identifying more instances of AI models violating established restrictions, including in the past when concerns about their irresponsible behavior were not as high as they are today. The incident involving a cyberattack on Hugging Face by an OpenAI model was a wake-up call for the industry, acknowledged Anthropic founder and CEO Dario Amodei. 'It’s easy to brush off this episode: no humans were harmed, and the economic damage was minimal. But such a swarm [of AI agents] with significant capabilities and equally serious disarray could cause catastrophic damage,' he wrote in an open letter, calling for a slowdown in AI development. Amodei made this statement after his employee Jacob Coxon left Anthropic, stating that new technologies 'could kill us all by the end of the decade.' Evan Hubinger, who leads research on managing and controlling future AI systems at the company, assessed the risk of human extinction due to AI actions in the next decade at over 10%, The Moscow Times reports.