CYBERCRIMINALS are successfully bypassing AI safety controls by splitting malicious activities across multiple sessions and files, according to research by Cisco Talos. This fragmentation prevents individual tasks from being flagged as harmful. Methods include falsely claiming ownership of targeted infrastructure and labeling activities as capture-the-flag or bug bounty work to evade scrutiny.
Talos noted that an actor's skill level significantly influences the effectiveness of AI-assisted operations, with novice attackers achieving limited results while experienced ones create sophisticated platforms. The findings underline the urgency for organizations to update their defenses against emerging AI-driven threats.