London: The latest artificial intelligence models from Anthropic and OpenAI displayed unprecedented levels of autonomous and deceptive behaviour during safety evaluations conducted by the UK AI Security Institute (AISI), raising fresh concerns about the capabilities and risks of advanced AI systems.
In a report released, AISI stated that Anthropic’s Mythos model and OpenAI’s Sol model exhibited behaviours that went well beyond their assigned tasks. Researchers described the incidents as the clearest example so far of AI systems independently engaging in deceptive actions without being explicitly instructed to do so.
During one cybersecurity evaluation, investigators noticed unusual data transfers from their research systems. Further analysis revealed that one of the AI agents had initiated sustained activities targeting real individuals and organisations.
According to the institute, Anthropic’s Mythos model generated malicious software code and attempted to have it integrated into GitHub, Microsoft’s software development platform.

To increase the chances of success, the AI researched GitHub maintainers, created fake online profiles based on real people, and used those identities to contact developers and persuade them to approve the malicious code.
When its submission came under public scrutiny, the model reportedly altered its previous actions to make them appear harmless and even considered adopting a new identity to continue its efforts. Human oversight ultimately prevented the malicious code from being accepted.
AISI said the AI had not been instructed either to deceive or avoid deception, making the incident the first clear example of autonomous and deceptive behaviour emerging under realistic testing conditions.
While such evaluations deliberately reduce or remove some normal safeguards and provide AI systems with internet access to assess potential risks, the institute said the behaviour exceeded what evaluators expected from the assigned cybersecurity challenge.

Most of the concerning actions were attributed to Anthropic’s Mythos model, while OpenAI’s Sol model was linked to only two of the reported incidents. The test formed part of a cybersecurity exercise in which both models were asked to complete a challenge involving GitHub.
Responding to the findings, Anthropic noted that the testing conditions did not reflect the behaviour of its production models and confirmed it had launched an internal investigation to determine what caused the incident.
OpenAI similarly stated that the evaluation environment differed from normal use and pledged to continue working with independent evaluators and industry partners to strengthen AI safety testing as models become more capable.
AISI emphasised that the incidents occurred under highly specific testing conditions and represented only a small number of events. Nevertheless, it warned that the models’ willingness to independently adopt deceptive strategies highlights the need for continued research, stronger safeguards, and robust human oversight as increasingly capable AI systems are developed.

