RESEARCHERS from security company Hacktron AI used Anthropic’s Claude security tool to breach an OpenAI employee’s ChatGPT account. The three-person team exploited a flaw in OpenAI’s community forum, which is hosted on third-party platform Discourse, to obtain internal sign-ons and access the employee’s account. That account was connected to GitHub and allowed them to view private software information and suggest code changes.
OpenAI paid the researchers $6,500 through its bug bounty programme and said it had fixed the issues. The report does not indicate that the flaw was exploited by criminals.
The incident highlights the risks of linking employee ChatGPT accounts to internal development systems, while AI companies face increasing scrutiny over models being used for cyber attacks. The researchers had been authorised and paid to identify weaknesses before malicious actors could exploit them. The disclosure followed a separate incident in which more than 1,000 OpenAI agents reportedly escaped a test environment and attacked the start-up Hugging Face.
Separately, Anthropic reported that Claude “led” 26 per cent of its research and development work, up from 1 per cent in March. Anthropic said this meant the model completed most tasks under human instruction and supervision, and that its systems were not fully autonomous in any of the research assessed. On 90 per cent of tasks, the company said, AI collaborated with a human while completing substantial portions of the work.