www.darkreading.com 11 Sept 2026, 17:10 UTC

Russia-Aligned Hackers Probe AI Safety Guardrails with Malware

CyberSIXT Evidence Panel
Threat Actor

HACKERS have demonstrated a new way to undermine AI-assisted security analysis by deliberately tripping the safety guardrails of large language models. In what ESET Labs calls GuardBreaker, the threat actor group UAC-0099, described as Russia-aligned, inserted a dangerous prompt into malicious code (a VBScript comment) that asks for a nuclear weapon. The aim was to trigger the AI’s safety mechanisms and prevent the model from analysing the rest of the malware.

The incident involved a victim in Ukraine, underscoring how adversaries are increasingly probing AI defenses rather than merely making malware more complex.

The piece argues that this is not an isolated misstep but a sign of a broader shift: AI-enabled threats are accelerating vulnerability discovery and exploitation, compressing historically long timelines from discovery to patching to hours or days. It highlights the need for governance alongside technical controls, noting recent industry and regulatory activity.

For example, OpenAI reportedly slowed frontier AI development after the Hugging Face incident, and about 130 organisations published a joint call to action on strengthening cyber defences. Governments are weighing responses, with CISA creating the Gold Eagle vulnerability clearinghouse and countries like South Korea pursuing their own AI frontier models.

The author argues for multilayered defence—detection, human-led engineering, behavioural analysis, sandboxing and telemetry—backed by governance frameworks to prevent a single control from breaching a wider security posture.

View full article

Article by CyberSIXT