Nation-StateHelp Net Security·4 hours ago
Russian hackers plant nuclear weapon prompt in malware to trip AI safety guardrails
Russian state-aligned hackers from UAC-0099 are embedding malicious prompts in malware to deliberately trigger AI safety guardrails and disrupt AI-assisted malware analysis tools, a technique ESET has dubbed GuardBreaker. The group, which conducts initial-access operations for the GRU-linked Sandworm, uses VBS scripts containing provocative prompts designed to overwhelm AI safety mechanisms and prevent proper analysis of their malicious code. This represents a new adversarial approach where threat actors actively work to subvert the security tools being used to detect their operations.