Adversarial Prompt Injection Attacks Help Cybercriminals Evade AI Security Tools

Cybercriminals are increasingly testing adversarial prompt injection attacks designed to bypass AI-powered security tools, according to recent threat research.

The technique targets large language model (LLM) applications and AI agents that scan emails, documents, calendar invitations, websites, and advertisements for malicious content.

AI has already changed the cyber threat landscape. Threat actors use LLM-assisted tools to create phishing lures, generate malicious code, improve social engineering, and scale campaigns faster.

Security teams also rely on AI to analyze suspicious content, prioritize alerts, and summarize emails or files. However, this growing use of AI creates another target for attackers, the AI security tool itself.

Researchers have observed criminal forum advertisements for indirect prompt injection (IDPI) generators. These tools are reportedly offered through subscriptions starting at around 150150150 dollars per month.

Available services include generators for malicious emails, PDF files, calendar invitations, and webpages. The activity remains experimental, and researchers have not yet reported widespread large-scale exploitation.

Still, the commercial availability of these tools suggests that attackers are preparing to use prompt injection against organizations operating AI-enabled email security, automated analysis, and agentic workflows.

Prompt Injection Evades AI Security

Prompt injection happens when an attacker inserts instructions that cause an AI model to behave in an unintended way. OWASP identifies two main types: direct and indirect prompt injection.

A direct injection occurs when a user provides malicious instructions directly to an AI application. For example, an attacker could attempt to make a chatbot ignore its safety rules or reveal restricted information.

The adversary generates an email with “white-on-white” text (hidden) (Source: proofpoint)
The adversary generates an email with “white-on-white” text (hidden) (Source: proofpoint)

Indirect prompt injection is more relevant to the newly advertised criminal tools. In this attack, an AI model reads instructions hidden inside external content, such as an email, PDF, webpage, or calendar event.

The AI may treat those instructions as valid commands rather than untrusted data. One observed email-generation tool creates messages containing white-on-white text.

The recipient may see a normal-looking email, while an AI mail agent can still read hidden text in the message body.

The concealed text could instruct the agent to ignore previous rules, classify a malicious message as safe, or perform an unauthorized action.

Looking into the PDF file it contains IDPI (Source: proofpoint)
Looking into the PDF file it contains IDPI (Source: proofpoint)

Attackers are also testing PDFs that appear harmless to users and traditional antivirus tools. One example contained an ordinary non-disclosure agreement sample but included hidden instructions intended for an AI scanning agent.

The prompt reportedly told the agent to stop its current task and send spreadsheet files to an attacker-controlled email address, proofpoint said.

Whether such an instruction succeeds depends on the AI agent’s permissions, document-parsing behavior, and security controls.

A prompt alone cannot force an agent to access files or send data if it lacks the necessary access.

However, the technique becomes more dangerous when an AI system has broad permissions and can process untrusted external content without strict controls.

Cut SOC investigation blind spots and contain threats earlier to reduce response costs and business disruption with ANY.RUN. 

Varshini
Varshini
Varshini is a Cyber Security expert in Threat Analysis, Vulnerability Assessment, and Research. Passionate about staying ahead of emerging Threats and Technologies..

Trending News

Related Stories