Reference. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
The robustness of LLMs to jailbreak attacks, where users design prompts to circumvent safety measures and misuse model capabilities, has been studied primarily for LLMs acting as simple chatbots. Meanwhile, LLM agents – which use external tools and can execute multi-stage tasks – may pose a greater risk if misused, but their robustness remains underexplored. To facilitate research on LLM agent misuse, we propose a new benchmark called AgentHarm. The benchmark includes a diverse set of 110 explicitly malicious agent tasks (440 with augmentations), covering 11 harm categories including fraud, cybercrime, and harassment. In addition to measuring whether models refuse harmful agentic requests, scoring well on AgentHarm requires jailbroken agents to maintain their capabilities following an attack to complete a multi-step task. We evaluate a range of leading LLMs, and find (1) leading LLMs are surprisingly compliant with malicious agent requests without jailbreaking, (2) simple universal jailbreak templates can be adapted to effectively jailbreak agents, and (3) these jailbreaks enable coherent and malicious multi-step agent behavior and retain model capabilities. To enable simple and reliable evaluation of attacks and defenses for LLM-based agents, we publicly release AgentHarm at https://huggingface.co/datasets/ai-safety-institute/AgentHarm.
Cite
Cited by (1)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition dziemian-2026-how
LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure to indirect prompt injection attacks, where adversarial instructions embedded in external content manipulate agent behavior without user awareness. A critical but underexplored dimension of this threat is concealment: since users tend to observe only an agent’s final response, an attack can conceal its existence by presenting no clue of compromise in the final user facing response while successfully executing harmful actions. This leaves users unaware of the manipulation and likely to accept harmful outcomes as legitimate. We present findings from a large scale public red teaming competition evaluating this dual objective across three agent settings: tool calling, coding, and computer use. The competition attracted 464 participants who submitted 272000 attack attempts against 13 frontier models, yielding 8648 successful attacks across 41 scenarios. All models proved vulnerable, with attack success rates ranging from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). We identify universal attack strategies that transfer across 21 of 41 behaviors and multiple model families, suggesting fundamental weaknesses in instruction following architectures. Capability and robustness showed weak correlation, with Gemini 2.5 Pro exhibiting both high capability and high vulnerability. To address benchmark saturation and obsoleteness, we will endeavor to deliver quarterly updates through continued red teaming competitions. We open source the competition environment for use in evaluations, along with 95 successful attacks against Qwen that did not transfer to any closed source model. We share model-specific attack data with respective frontier labs and the full dataset with the UK AISI and US CAISI to support robustness research.
Cites 38 works (0 here)
External (38)
- Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities (2024)
- Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks (2024)
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet (2024)
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models (2024)
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities (2024)
- Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification (2024)
- The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models (2024)
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents (2024)
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases (2024)
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents (2024)
- GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning (2024)
- Improving Alignment and Robustness with Circuit Breakers (2024)
- BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards (2024)
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks (2024)
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models (2024)
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents (2024)
- PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails (2024)
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024)
- Many-shot Jailbreaking (2024)
- A StrongREJECT for Empty Jailbreaks (2024)
- Low-Resource Languages Jailbreak GPT-4 (2024)
- Berkeley Function Calling Leaderboard (2024)
- Inspect AI: Framework for Large Language Model Evaluations (2024)
- Exploiting Novel GPT-4 APIs (2023)
- Evaluating Language-Model Agents on Realistic Autonomous Tasks (2023)
- GAIA: a benchmark for General AI Assistants (2023)
- Evil Geniuses: Delving into the Safety of LLM-based Agents (2023)
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox (2023)
- Calibrating LLM-Based Evaluator (2023)
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs (2023)
- Universal and Transferable Adversarial Attacks on Aligned Language Models (2023)
- Judging LLM-as-a-judge with MT-Bench and Chatbot Arena (2023)
- Gorilla: Large Language Model Connected with Massive APIs (2023)
- Emergent autonomous scientific research capabilities of large language models (2023)
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark (2023)
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
- ChemCrow: Augmenting large-language models with chemistry tools (2023)
- ReAct: Synergizing Reasoning and Acting in Language Models (2022)