Reference. Boundary Point Jailbreaking of Black-Box LLMs
Frontier LLMs are safeguarded against attempts to extract harmful information via adversarial prompts known as “jailbreaks”. Recently, defenders have developed classifier-based systems that have survived thousands of hours of human red teaming. We introduce Boundary Point Jailbreaking (BPJ), a new class of automated jailbreak attacks that evade the strongest industry-deployed safeguards. Unlike previous attacks that rely on white/grey-box assumptions (such as classifier scores or gradients) or libraries of existing jailbreaks, BPJ is fully black-box and uses only a single bit of information per query: whether or not the classifier flags the interaction. To achieve this, BPJ addresses the core difficulty in optimising attacks against robust real-world defences: evaluating whether a proposed modification to an attack is an improvement. Instead of directly trying to learn an attack for a target harmful string, BPJ converts the string into a curriculum of intermediate attack targets and then actively selects evaluation points that best detect small changes in attack strength (“boundary points”). We believe BPJ is the first fully automated attack algorithm that succeeds in developing universal jailbreaks against Constitutional Classifiers, as well as the first automated attack algorithm that succeeds against GPT-5′s input classifier without relying on human attack seeds. BPJ is difficult to defend against in individual interactions but incurs many flags during optimisation, suggesting that effective defence requires supplementing single-interaction methods with batch-level monitoring.
Cite
Cites 40 works (0 here)
External (40)
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks (2026)
- Black-box Optimization of LLM Outputs by Asking for Directions (2025)
- STACK: Adversarial Attacks on LLM Safeguard Pipelines (2025)
- Universal Jailbreak Suffixes Are Strong Attention Hijackers (2025)
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming (2025)
- Stronger Universal and Transferable Attacks by Suppressing Refusals (2025)
- Introducing GPT-4.1 in the API (2025)
- From bugs to bypasses: Adapting vulnerability disclosure for AI safeguards (2025)
- Prompt injection is not SQL injection (it may be worse) (2025)
- Automatically jailbreaking frontier language models with investigator agents (2025)
- The attacker moves second: Stronger adaptive attacks bypass defenses against LLM jailbreaks and prompt injections (2025)
- Strengthening our safeguards through collaboration with US CAISI and UK AISI (Anthropic) (2025)
- Constitutional classifiers: Defending against universal jailbreaks (Anthropic research post) (2025)
- Working with US CAISI and UK AISI to build more secure AI systems (OpenAI) (2025)
- GPT-5 System Card (2025)
- Security challenges in AI agent deployment: Insights from a large scale public competition (2025)
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks (2024)
- Fast Adversarial Attacks on Language Models In One GPU Minute (2024)
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024)
- HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text (2024)
- Expanding our model safety bug bounty program (2024)
- "Do Anything Now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models (2024)
- Query-based adversarial prompt generation (2024)
- Best-of-N jailbreaking (2024)
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (2023)
- Jailbreaking Black Box Large Language Models in Twenty Queries (2023)
- Universal and Transferable Adversarial Attacks on Aligned Language Models (2023)
- LeapAttack: Hard-Label Adversarial Attack on Text via Gradient-Based Optimization (2022)
- HopSkipJumpAttack: A Query-Efficient Decision-Based Attack (2019)
- Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach (2018)
- Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models (2017)
- Delving into Transferable Adversarial Examples and Black-box Attacks (2016)
- Curriculum learning (2009)
- Active Learning Literature Survey (2009)
- Numerical Continuation Methods for Dynamical Systems: Path following and boundary value problems (2007)
- Query by committee (1992)
- Some Guidelines and Guarantees for Common Random Numbers (1992)
- Adaptation in Natural and Artificial Systems: An Introductory Analysis with Applications to Biology, Control, and Artificial Intelligence (1992)
- Numerical Continuation Methods: An Introduction (1990)
- Selection and Covariance (1970)