Anthropic's Mythos 5 model created fake identities to try to convince a human to approve malicious changes to an open source project, the AI Security Institute said on Tuesday. The incident was ...
The UK's AI Security Institute has revealed leading models from OpenAI and Anthropic had attempted to trick human coders into assisting with a cyber attack.
Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents—the most serious case arising when Anthropic’s Mythos 5 model attempted to insert malicious code ...