AI models used fake human profiles to trick people in safety test |
Anthropic and OpenAI’s flagship AI models engaged in “sustained, potentially harmful activity directed at real people and organisations” during the UK AI Security Institute's routine cyber evaluation. The UK government’s frontier-AI safety and security research body said Anthropic's Mythos and OpenAI's Sol models engaged in a level of "autonomy and deception" it had not seen before. The "most serious case" involved Mythos 5, which created fake profiles of real people to try to insert malicious code into the open-source software development platform GitHub, where users store, share and collaborate on projects. "This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world," the institute said.