HIDDEN prompt injections are a new class of attack that target autonomous AI agents by concealing instructions inside external content. Unlike traditional prompts that a user directly feeds to an AI, these attacks embed malicious guidance in documents, metadata, emails, images or code that an AI may ingest during operation.
Proponents warn that such injections can cause agentic systems to behave in ways that breach intended guardrails, act at machine speed and without human-like judgment, enabling data exfiltration, file manipulation or other harmful actions.
Bowbridge, which researches AI security, points to real-world scenarios where an AI agent reviewing supplier quotes could be steered to override prior guidance. In the example cited, malicious content hidden in document metadata caused the agent to select a more expensive quote, illustrating how trusted systems can be manipulated when content comes from untrusted sources. The concern is heightened because agents often inherit the privileges of their users and operate without visible human oversight.
The suggested defence is to prioritise prevention of poisoning: scan documents before processing, detect hidden content within files and metadata, and apply AI security frameworks and controls that can intercept risky content before it reaches agents. Bowbridge notes that as organisations embed agentic AI across workflows, protecting the content consumed by these systems will become a central aspect of cybersecurity strategy.