securityonline.info 14 Sept 2026, 08:01 UTC

Puzzlemask Tricks AI Gatekeepers Into Executing Hidden Attacks

Puzzlemask Tricks AI Gatekeepers Into Executing Hidden Attacks
CyberSIXT Evidence Panel Source marked as original reporting

CHECK Point Research has documented “Puzzlemask”, a prompt-injection technique that can bypass lightweight language-model security checks without using emojis, Base64, invisible Unicode or other conspicuous encoding. Instead, an attacker embeds instructions in apparently ordinary English prose. In multi-tier AI systems, a smaller gatekeeper model may classify the text as harmless and pass it to a more capable model, which can identify the embedded structure and carry out the payload. The researchers said the technique could be used to request actions including file encryption, data theft or system compromise.

In testing, Check Point Research submitted 23 crafted prompts to GPT-4o mini, Claude 3 Haiku, Llama-Guard 3 and gpt-oss-safeguard. The gatekeeper models classified the crafted prompts as safe in 100% of the trials. Eighteen prompts were then sent to GPT-5 Thinking High with access to an internal Python interpreter; it decoded and executed the hidden instructions in 17 cases, or 94.4%. In one emulated example, the model encrypted local documents.

The attack required more than a minute of reasoning and several Python scripts, while Anthropic’s Opus-class systems reportedly stopped when they detected obfuscated prompts.

The researchers recommend paraphrasing untrusted input before processing, although this may add delay and change meaning. They also found that explicitly instructing gatekeepers to flag quantitative, self-referential statements enabled GPT-4o mini to intercept every test prompt. Monitoring tool calls and resulting actions provides an additional safeguard beyond input screening.

View full article

Article by CyberSIXT

Timeline Coverage

Swipe to explore timeline