RECENT research from Forcepoint X-Labs reveals how attackers can manipulate AI email summarizers with hidden prompts within emails, leading to the generation of false summaries. This vulnerability arises from AI's struggle to differentiate between data and instructions during processing. In a controlled experiment, researchers created emails with malicious hidden HTML prompts, successfully altering AI-generated summaries in multiple trials.
The findings underscore the risk of 'prompt injection' attacks, highlighting the need for organizations to treat incoming content and AI outputs as untrusted, implement safeguards, and separate user-generated content from instructions given to AI models.