AT Black Hat USA 2026, Simcha Kosman from Palo Alto Networks presented a proof-of-concept attack detailing how a malicious actor could gain command and control (C2) over an isolated ChatGPT sandbox. The attack exploits differences in URL handling between platforms, allowing attackers to execute commands automatically on certain devices.
By manipulating how ChatGPT processes files, including running embedded code in spreadsheets and using shared resources like JFrog's Artifactory, attacks could send sensitive data back to the attacker's server. Despite being theoretical, the research highlights vulnerabilities within sandbox systems, prompting OpenAI to implement security changes following the findings.