A hacker known as "Trim" has transformed publicly available large language models (LLMs) into a commercial offensive cybercriminal tool after freely publishing jailbreaking techniques online. Trim's methods, outlined on a Russian-language forum, include six techniques that manipulate AI models to bypass safety filters. These techniques include using legitimate prompts to build trust, switching to different AI models when faced with refusals, and employing low-cost API keys from underground sources.
After initially sharing his methods, Trim launched the "AI Pentest Checker," a service that automates web vulnerability scanning using AI tools. Experts warn that such developments in AI technology lower barriers for cybercriminals, making powerful tools accessible to a wider range of actors. This evolution in offensive capabilities underscores the urgency for cybersecurity defenders to adapt their strategies, emphasizing the need for increased visibility and monitoring of potential vulnerabilities.