Reference. Identifying and Mitigating the Security Risks of Generative AI
Every major technical invention resurfaces the dual-use dilemma—the new technology has the potential to be used for good as well as for harm. Generative AI (GenAI) techniques, such as large language models (LLMs) and diffusion models, have shown remarkable capabilities (e.g., in-context learning, code-completion, and text-to-image generation and editing). However, GenAI can be used just as well by attackers to generate new attacks and increase the velocity and efficacy of existing attacks. This monograph reports the findings of a workshop held at Google (co-organized by Stanford University and the University of Wisconsin-Madison) on the dual-use dilemma posed by GenAI. This work is not meant to be comprehensive, but is rather an attempt to synthesize some of the interesting findings from the workshop. We discuss short-term and long-term goals for the community on this topic. We hope this work provides both a launching point for a discussion on this important topic as well as interesting problems that the research community can work to address.
Cite
Cites 108 works (0 here)
External (108)
- Poisoning Web-Scale Training Datasets is Practical (2024)
- Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks (2024)
- Towards the Detection of Diffusion Model Deepfakes (2024)
- Product liability for defective AI (2024)
- The Stable Signature: Rooting Watermarks in Latent Diffusion Models (2023)
- RARR: Researching and Revising What Language Models Say, Using Language Models (2023)
- Large Language Models for Code: Security Hardening and Adversarial Testing (2023)
- Evading Watermark based Detection of AI-Generated Content (2023)
- GPT detectors are biased against non-native English writers (2023)
- PTW: Pivotal tuning watermarking for pre-trained image generators (2023)
- Towards universal fake image detectors that generalize across generative models (2023)
- In-Context Retrieval-Augmented Language Models (2023)
- Measuring Attribution in Natural Language Generation Models (2023)
- DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation Models (2023)
- ChatGPT: More than a ‘weapon of mass deception’ ethical challenges and responses from the human-centered artificial intelligence (HCAI) perspective (2023)
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts (2023)
- Can Large Language Models Transform Computational Social Science? (2023)
- ReAct: Synergizing Reasoning and Acting in Language Models (2023)
- Self-Refine: Iterative Refinement with Self-Feedback (2023)
- Universal and Transferable Adversarial Attacks on Aligned Language Models (2023)
- Can AI-Generated Text be Reliably Detected? (2023)
- The Curse of Recursion: Training on Generated Data Makes Models Forget (2023)
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature (2023)
- ChemCrow: Augmenting large-language models with chemistry tools (2023)
- A Watermark for Large Language Models (2023)
- On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective (2023)
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense (2023)
- GPT-4 is OpenAI's most advanced system, producing safer and more useful responses (web page) (2023)
- Are aligned neural networks adversarially aligned? (2023)
- On the Reliability of Watermarks for Large Language Models (2023)
- LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? (2023)
- Robust Distortion-free Watermarks for Language Models (2023)
- Provable Robust Watermarking for AI-Generated Text (2023)
- Detecting Language Model Attacks with Perplexity (2023)
- Undetectable Watermarks for Language Models (2023)
- Rethinking Model Evaluation as Narrowing the Socio-Technical Gap (2023)
- Authors Guild letter seeks compensation from AI companies for using authors' writings in AI (2023)
- ChaosGPT: Empowering GPT with Internet and Memory to Destroy Humanity (video) (2023)
- Securing the Future of GenAI: Mitigating Security Risks (workshop site) (2023)
- OpenAI, Google will watermark AI-generated content to hinder deepfakes, misinfo (2023)
- Lawyer Used ChatGPT In Court - And Cited Fake Cases. A Judge Is Considering Sanctions (2023)
- Lawsuit says OpenAI violated US authors' copyrights to train AI chatbot (2023)
- Tuning Models of Code with Compiler-Generated Reinforcement Learning Feedback (2023)
- WormGPT - The Generative AI Tool Cybercriminals Are Using to Launch Business Email Compromise Attacks (2023)
- FraudGPT: The Villain Avatar of ChatGPT (2023)
- Introducing Llama 2 (Meta) (2023)
- Stable Diffusion (website) (2023)
- Long Sequence Modeling with XGen: A 7B LLM Trained on 8K Input Sequence Length (2023)
- DALL-E: Creating images from text (OpenAI) (2023)
- GPT-4 Technical Report (2023)
- My AI: Snapchat chatbot coaches 'girl, 13' on losing virginity (2023)
- Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence (2023)
- FACT SHEET: Biden-Harris Administration Announces National Cyber Workforce and Education Strategy (2023)
- Detecting child sexual abuse material shouldn't be done at any cost (2023)
- Dual-use technology - Wikipedia (2023)
- Large language model - Wikipedia (2023)
- 'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says (2023)
- Protecting Language Generation Models via Invisible Watermarking (2023)
- TRUE: Re-evaluating Factual Consistency Evaluation (2022)
- Training language models to follow instructions with human feedback (2022)
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions (2022)
- Red teaming language models with language models (2022)
- Emergent Abilities of Large Language Models (2022)
- Teaching language models to support answers with verified quotes (2022)
- Data Feedback Loops: Model-driven Amplification of Dataset Biases (2022)
- Constitutional AI: Harmlessness from AI Feedback (2022)
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned (2022)
- Poisoning and Backdooring Contrastive Learning (2022)
- How the EU can take on the challenge posed by general-purpose AI systems (2022)
- Temporary policy: Generative AI (e.g., ChatGPT) is banned (Stack Overflow) (2022)
- ChatGPT: Optimizing Language Models for Dialogue (OpenAI blog) (2022)
- TweepFake: About detecting deepfake tweets (2021)
- The importance of modeling social factors of language: Theory and practice (2021)
- Challenges in Detoxifying Language Models (2021)
- Process for Adapting Language Models to Society (PALMS) with\n Values-Targeted Datasets (2021)
- DARPA Announces Research Teams Selected to Semantic Forensics Program (2021)
- Generating sentiment-preserving fake online reviews using neural language models and their human-and machine-based detection (2020)
- Extracting training data from large language models (2020)
- Leveraging frequency analysis for deep fake image recognition (2020)
- Artificial intelligence, values, and alignment (2020)
- RealToxicityPrompts: Evaluating neural toxic degeneration in language models (2020)
- Deepfake detection by analyzing convolutional traces (2020)
- Automatic Detection of Generated Text is Easiest when Humans are Fooled (2020)
- Automatic Detection of Machine Generated Text: A Critical Survey (2020)
- Global texture enhancement for fake face detection in the wild (2020)
- CNN-generated images are surprisingly easy to spot for now (2020)
- Recipes for Safety in Open-domain Chatbots (2020)
- Learning to summarize with human feedback (2020)
- GLTR: Statistical Detection and Visualization of Generated Text (2019)
- Incremental learning for the detection and classification of GAN-generated images (2019)
- Detecting GAN generated Fake Images using Co-occurrence Matrices (2019)
- Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions (2019)
- Detecting and simulating artifacts in gan fake images (2019)
- Detecting GAN-Generated Imagery Using Saturation Cues (2019)
- BERTScore: Evaluating Text Generation with BERT (2019)
- Release Strategies and the Social Impacts of Language Models (2019)
- Unmasking DeepFakes with simple Features (2019)
- Real or fake? Learning to discriminate machine from human generated text (2019)
- RoBERTa: A Robustly Optimized BERT Pretraining Approach (2019)
- GPT-2: 1.5B release (OpenAI) (2019)
- Detection of gan-generated fake images over social networks (2018)
- Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints (2017)
- Attention is All you Need (2017)
- Information Hiding (Katzenbeisser & Petitcolas) (2016)
- Linguistic steganography on Twitter: Hierarchical language modeling with manual interaction (2014)
- Natural language watermarking: Design, analysis, and a proof-of-concept implementation (2001)
- The intellectual challenge of CSCW: The gap between social requirements and technical feasibility (2000)
- Your AI model might be telling you this is not a cat (web demo)