Reference. Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
LLM developers have imposed technical interventions to prevent fine-tuning misuse attacks, attacks where adversaries evade safeguards by fine-tuning the model using a public API. Previous work has established several successful attacks against specific fine-tuning API defences. In this work, we show that defences of fine-tuning APIs that seek to detect individual harmful training or inference samples (‘pointwise’ detection) are fundamentally limited in their ability to prevent fine-tuning attacks. We construct ‘pointwise-undetectable’ attacks that repurpose entropy in benign model outputs (e.g. semantic or syntactic variations) to covertly transmit dangerous knowledge. Our attacks are composed solely of unsuspicious benign samples that can be collected from the model before fine-tuning, meaning training and inference samples are all individually benign and low-perplexity. We test our attacks against the OpenAI fine-tuning API, finding they succeed in eliciting answers to harmful multiple-choice questions, and that they evade an enhanced monitoring system we design that successfully detects other fine-tuning attacks. We encourage the community to develop defences that tackle the fundamental limitations we uncover in pointwise fine-tuning API defences.
Cite
Cites 47 works (0 here)
External (47)
- Benchmarking Misuse Mitigation Against Covert Adversaries (2025)
- Managing Misuse Risk for Dual-Use Foundation Models (NIST AI 800-1, 2nd public draft) (2025)
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025) (2025)
- OpenAI Fine-tuning guide (2025)
- Deliberative Alignment: Reasoning Enables Safer Language Models (2024)
- Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats (2024)
- Measuring short-form factuality in large language models (2024)
- FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs? (2024)
- Tamper-Resistant Safeguards for Open-Weight LLMs (2024)
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs (2024)
- Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries (2024)
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep (2024)
- OR-Bench: An Over-Refusal Benchmark for Large Language Models (2024)
- Representation Noising: A Defence Mechanism Against Harmful Finetuning (2024)
- What is in Your Safe Data? Identifying Benign Data that Breaks Safety (2024)
- Immunization against harmful fine-tuning attacks (2024)
- Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment (2024)
- Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack (2024)
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024)
- Secret Collusion Among Generative AI Agents (2024)
- Fine-tune anthropic’s claude 3 haiku in amazon bedrock to boost model accuracy and quality (2024)
- Fine-tuning with the Gemini API (2024)
- Customize a model with fine-tuning: Safety evaluation GPT-4, GPT-4o, and GPT-4o-mini fine-tuning - public preview (2024)
- Generative AI Prohibited Use Policy (2024)
- OpenAI o1 System Card (2024)
- Anthropic Usage Policy (2024)
- Sabotage evaluations for frontier models (2024)
- Poisoning web-scale training datasets is practical (2024)
- Content filtering (Azure OpenAI documentation) (2024)
- Covert malicious finetuning: Challenges in safeguarding LLM adaptation (2024)
- The WMDP benchmark: Measuring and reducing malicious use with unlearning (2024)
- Llama Guard 3 documentation (2024)
- OpenAI Model Specification - May 8, 2024 (2024)
- OpenAI Usage Policies (2024)
- Exploiting Novel GPT-4 APIs (2023)
- Look Before You Leap: A Universal Emergent Decomposition of Retrieval Tasks in Language Models (2023)
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks (2023)
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023)
- Removing RLHF Protections in GPT-4 via Fine-Tuning (2023)
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (2023)
- Instruction Tuning with GPT-4 (2023)
- Auditing failures vs concentrated failures (2023)
- Measuring faithfulness in chain-of-thought reasoning (2023)
- Low-stakes alignment (2021)
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge (2019)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning (2017)
- Principles and practice of information theory (1987)