More than you've asked for: A Comprehensive Analysis of Novel Prompt Injection Threats to Application-Integrated Large Language Models
Explore this paper's citation graph
Summary
This work shows that augmenting LLMs with retrieval and API calling capabilities (so-called Application-Integrated LLMs) induces a whole new set of attack vectors and systematically analyzes the resulting threat landscape of Application-Integrated LLMs.
- Type
- preprint
- Published
- 2023-02-23
- Cited by
- 157
- References
- 37
- Access
- Open access
- OpenAlex
- https://openalex.org/W4321855128
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:257102404
Keywords
Exploit, Computer science, Computer security, Interface (matter), Software deployment
References
- StereoSet: Measuring stereotypical bias in pretrained language models
- It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- “Was it “stated” or was it “claimed”?: How linguistic bias affects generative language models
- Get a Model! Model Hijacking Attack Against Machine Learning Models
- Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
- Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations
- Ignore Previous Prompt: Attack Techniques For Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- “Real Attackers Don't Compute Gradients”: Bridging the Gap Between Adversarial ML Research and Practice
- Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
- Training language models to follow instructions with human feedback
- Chain of Thought Prompting Elicits Reasoning in Large Language Models
- Understanding Emails and Drafting Responses - An Approach Using GPT-3
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- ReAct: Synergizing Reasoning and Acting in Language Models
- Language Models are Few-Shot Learners
- Ethical and social risks of harm from Language Models
Cited by
- Measuring and Manipulating Knowledge Representations in Language Models
- Multi-step Jailbreaking Privacy Attacks on ChatGPT
- In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT
- Appropriateness is all you need!
- Tricking LLMs into Disobedience: Understanding, Analyzing, and Preventing Jailbreaks
- Responsible Task Automation: Empowering Large Language Models as Responsible Task Automators
- Spear or Shield: Leveraging Generative AI to Tackle Security Threats of Intelligent Network Services
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- Trapping LLM Hallucinations Using Tagged Context Prompts
- Safeguarding Crowdsourcing Surveys from ChatGPT with Prompt Injection
- TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models
- Are aligned neural networks adversarially aligned?
- Evaluating GPT-3.5 and GPT-4 on Grammatical Error Correction for Brazilian Portuguese
- Code Injection Attacks in Wireless-Based Internet of Things (IoT): A Comprehensive Review and Practical Implementations
- Stay on topic with Classifier-Free Guidance
- Jailbroken: How Does LLM Safety Training Fail?
- Foundational Models Defining a New Era in Vision: A Survey and Outlook
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
- Mondrian: Prompt Abstraction Attack Against Large Language Models for Cheaper API Pricing
- Dual Governance: The intersection of centralized regulation and crowdsourced safety mechanisms for Generative AI
Related papers
- AEMB: An Automated Exploit Mitigation Bypassing Solution
- AEG: Automatic Exploit Generation
- Evaluation of Two Host-Based Intrusion Prevention Systems
- Exploit Kits: The production line of the Cybercrime economy?
- Automated Crash Analysis and Exploit Generation with Extendable Exploit Model
- EBF: Event-Based Filter for Exploit Containment
- THE SEARCH ON THE EXPLOIT MACNINATION OF THEME PARK