OPENAI has introduced new security protocols for its AI models, implementing stricter containment and monitoring measures in response to recent incidents highlighting vulnerabilities. The updated system includes enhanced sandboxes for executing untrusted code and a multistage monitoring framework that analyzes model activity for anomalies. Alerts are managed under a tight operational SLA, requiring immediate investigation or activity pauses if false positives are not ruled out in 30 minutes.
These measures will apply to models with high cybersecurity capabilities, reflecting the company’s goal to ensure AI models can effectively manage security operations and prevent future breaches. Similar security concerns have arisen with models from other companies, including Anthropic and Meta.