Reference. How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure to indirect prompt injection attacks, where adversarial instructions embedded in external content manipulate agent behavior without user awareness. A critical but underexplored dimension of this threat is concealment: since users tend to observe only an agent’s final response, an attack can conceal its existence by presenting no clue of compromise in the final user facing response while successfully executing harmful actions. This leaves users unaware of the manipulation and likely to accept harmful outcomes as legitimate. We present findings from a large scale public red teaming competition evaluating this dual objective across three agent settings: tool calling, coding, and computer use. The competition attracted 464 participants who submitted 272000 attack attempts against 13 frontier models, yielding 8648 successful attacks across 41 scenarios. All models proved vulnerable, with attack success rates ranging from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). We identify universal attack strategies that transfer across 21 of 41 behaviors and multiple model families, suggesting fundamental weaknesses in instruction following architectures. Capability and robustness showed weak correlation, with Gemini 2.5 Pro exhibiting both high capability and high vulnerability. To address benchmark saturation and obsoleteness, we will endeavor to deliver quarterly updates through continued red teaming competitions. We open source the competition environment for use in evaluations, along with 95 successful attacks against Qwen that did not transfer to any closed source model. We share model-specific attack data with respective frontier labs and the full dataset with the UK AISI and US CAISI to support robustness research.
Cite
Cites 66 works (1 here)
With notes (1)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents andriushchenko-2024-agentharm
The robustness of LLMs to jailbreak attacks, where users design prompts to circumvent safety measures and misuse model capabilities, has been studied primarily for LLMs acting as simple chatbots. Meanwhile, LLM agents – which use external tools and can execute multi-stage tasks – may pose a greater risk if misused, but their robustness remains underexplored. To facilitate research on LLM agent misuse, we propose a new benchmark called AgentHarm. The benchmark includes a diverse set of 110 explicitly malicious agent tasks (440 with augmentations), covering 11 harm categories including fraud, cybercrime, and harassment. In addition to measuring whether models refuse harmful agentic requests, scoring well on AgentHarm requires jailbroken agents to maintain their capabilities following an attack to complete a multi-step task. We evaluate a range of leading LLMs, and find (1) leading LLMs are surprisingly compliant with malicious agent requests without jailbreaking, (2) simple universal jailbreak templates can be adapted to effectively jailbreak agents, and (3) these jailbreaks enable coherent and malicious multi-step agent behavior and retain model capabilities. To enable simple and reliable evaluation of attacks and defenses for LLM-based agents, we publicly release AgentHarm at https://huggingface.co/datasets/ai-safety-institute/AgentHarm.
External (65)
- Building Production-Ready Probes For Gemini (2026)
- CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents (2026)
- Systems security foundations for agentic computing (2026)
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing (2025)
- Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents (2025)
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections (2025)
- "Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors (2025)
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models (2025)
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025)
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents (2025)
- Design Patterns for Securing LLM Agents against Prompt Injections (2025)
- Transferable Adversarial Attacks on Black-Box Vision-Language Models (2025)
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks (2025)
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection (2025)
- Defeating Prompt Injections by Design (2025)
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation (2025)
- Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks (2025)
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming (2025)
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-Based Prompt Injection Attacks via the Fine-Tuning Interface (2025)
- Amazon Nova 2: Multimodal reasoning and generation models (2025)
- Advancing Claude for financial services (2025)
- Claude Opus 4.5 system card (2025)
- Shade: Automated AI red-teaming (2025)
- cellmate: Sandboxing browser AI agents (2025)
- Cybench: A framework for evaluating cybersecurity capabilities and risks of language models (2025)
- Security challenges in AI agent deployment: Insights from a large scale public competition (2025)
- Deliberative Alignment: Reasoning Enables Safer Language Models (2024)
- Agent-SafetyBench: Evaluating the Safety of LLM Agents (2024)
- Best-of-N Jailbreaking (2024)
- Imprompter: Tricking LLM Agents into Improper Tool Use (2024)
- SecAlign: Defending Against Prompt Injection with Preference Optimization (2024)
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents (2024)
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions (2024)
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models (2024)
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents (2024)
- Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications (2024)
- StruQ: Defending Against Prompt Injection with Structured Queries (2024)
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024)
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs (2024)
- Cursor: The AI code editor (2024)
- GPQA diamond benchmark (2024)
- GitHub Copilot: Your AI pair programmer (2024)
- Hello GPT-4o (2024)
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations (2023)
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (2023)
- FigStep: Jailbreaking Large Vision-language Models via Typographic Visual Prompts (2023)
- Jailbreaking Black Box Large Language Models in Twenty Queries (2023)
- Misusing Tools in Large Language Models With Visual Adversarial Examples (2023)
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models (2023)
- Image Hijacks: Adversarial Images can Control Generative Models at Runtime (2023)
- Universal and Transferable Adversarial Attacks on Aligned Language Models (2023)
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
- Ignore Previous Prompt: Attack Techniques For Language Models (2022)
- MPNet: Masked and Permuted Pre-training for Language Understanding (2020)
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (2019)
- The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover
- Generative red team 3 (GRT3)
- Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for 19
- Devin AI Kill Chain—Exposing Ports Leading to RCE and file Exfiltration
- The state of AI in 2025: Agents, innovation, and transformation
- Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku
- Measuring ai ability to complete long tasks
- ChatGPT Operator prompt injection exploits
- Microsoft Copilot: From Prompt Injection to Exfil-tration of Personal Information
- Spyware Injection Into Your ChatGPT’s Long-Term Memory (SpAIware)