Automatically Auditing Large Language Models via Discrete Optimization
Explore this paper's citation graph
Summary
This work cast auditing as an optimization problem, where it automatically search for input-output pairs that match a desired target behavior, and introduces a discrete optimization algorithm, ARCA, that jointly and efficiently optimizes over inputs and outputs.
- Type
- preprint
- Published
- 2023-03-08
- Cited by
- 242
- References
- 123
- Access
- Open access
- OpenAlex
- https://openalex.org/W4323709479
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:257405439
Keywords
Audit, Software deployment, Computer science, Set (abstract data type), Space (punctuation)
References
- Intriguing properties of neural networks
- Bag of Tricks for Efficient Text Classification
- FastText.zip: Compressing text classification models
- HotFlip: White-Box Adversarial Examples for Text Classification
- Conversational AI: The Science Behind the Alexa Prize
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification
- Unrestricted Adversarial Examples
- Counterfactual Fairness in Text Classification through Robustness
- Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
- Categorical Reparameterization with Gumbel-Softmax
- Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products
- Generating Natural Language Adversarial Examples
- Adversarial Examples for Evaluating Reading Comprehension Systems
- Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
- Universal Adversarial Triggers for Attacking and Analyzing NLP
- The Woman Worked as a Babysitter: On Biases in Language Generation
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Plug and Play Language Models: A Simple Approach to Controlled Text Generation
- Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing
- BERT-ATTACK: Adversarial Attack against BERT Using BERT
Cited by
- Model evaluation for extreme risks
- Explore, Establish, Exploit: Red Teaming Language Models from Scratch
- Mass-Producing Failures of Multimodal Systems with Language Models
- Visual Adversarial Examples Jailbreak Aligned Large Language Models
- Are aligned neural networks adversarially aligned?
- Jailbroken: How Does LLM Safety Training Fail?
- On the Trustworthiness Landscape of State-of-the-art Generative Models: A Comprehensive Survey
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
- Detecting Language Model Attacks with Perplexity
- Image Hijacks: Adversarial Images can Control Generative Models at Runtime
- Large Language Model Alignment: A Survey
- Low-Resource Languages Jailbreak GPT-4
- LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model
- Jailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- Diversity of Thought Improves Reasoning Abilities of Large Language Models
- Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
- Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases
- AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
- Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Related papers
- Serving Away from Home: How Deployments Influence Reenlistment
- Understanding deployment from the perspective of those who have served.
- Research on Accelerating Application Technology of Centralized ERP System Based on HANA
- NATIONAL ITS PROGRAM PLAN : SYNOPSIS
- Heuristic algorithms for effective broker deployment