SENTINELONE introduced the first reverse-engineering benchmark for AI models, testing their ability to analyze the Fast16 malware linked to US-Iran cyber tensions. The study evaluated models like OpenAI's GPT-5.6 and others through eight stages of investigation, with only GPT-5.6 successfully completing all stages. Researchers noted that human oversight is crucial, as even the best ai models made significant errors and premature conclusions. This highlights the ongoing need for human analysts in cybersecurity roles.
Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models
CyberSIXT Evidence Panel
Primary Source
sentinelone.com
Article by CyberSIXT