BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Explore this paper's citation graph
Summary
It is shown that outsourced training introduces new security risks: an adversary can create a maliciously trained network (a backdoored neural network, or a BadNet) that has state-of-the-art performance on the user's training and validation samples, but behaves badly on specific attacker-chosen inputs.
- Type
- preprint
- Published
- 2017-08-22
- Cited by
- 2,382
- References
- 52
- Access
- Open access
- OpenAlex
- https://openalex.org/W2748789698
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:26783139
Keywords
Backdoor, Computer science, Debugging, Classifier (UML), Traffic sign recognition
References
- Good Word Attacks on Statistical Spam Filters
- Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach
- Learning algorithms for classification: A comparison on handwritten digit recognition
- On Attacking Statistical Spam Filters
- Original Contribution: Training a 3-node neural network is NP-complete
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Intriguing properties of neural networks
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Playing Atari with Deep Reinforcement Learning
- Traffic sign detection for U.S. roads: Remaining challenges and a case for tracking
- Correlating fourier descriptors of local patches for road sign recognition
- Deep convolutional neural network based species recognition for wild animal monitoring
- CNN Features Off-the-Shelf: An Astounding Baseline for Recognition
- Deep learning in neural networks: An overview
- Adversarial machine learning
- DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving
- Speech recognition with deep recurrent neural networks
- ImageNet classification with deep convolutional neural networks
- A Survey on Transfer Learning
- Rethinking the Inception Architecture for Computer Vision
Cited by
- MTDeep: Boosting the Security of Deep Neural Nets Against Adversarial Attacks with Moving Target Defense
- Hardening quantum machine learning against adversaries
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- PoTrojan: powerful neural-level trojan designs in deep learning models
- The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
- Transpositional neurocryptography using Deep Learning
- When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning Attacks
- Multi-source data analysis and evaluation of machine learning techniques for SQL injection detection
- Protecting Intellectual Property of Deep Neural Networks with Watermarking
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- How To Backdoor Federated Learning
- A Survey of Adversarial Machine Learning in Cyber Warfare
- Security and Privacy Issues in Deep Learning
- When Deep Learning Meets Inter-Datacenter Optical Network Management: Advantages and Vulnerabilities
- Mitigating Sybils in Federated Learning Poisoning
- VerIDeep: Verifying Integrity of Deep Neural Networks through Sensitive-Sample Fingerprinting
- Using Randomness to Improve Robustness of Machine-Learning Models Against Evasion Attacks
- Have You Stolen My Model? Evasion Attacks Against Deep Neural Network Watermarking Techniques
Related papers
- I Know Your Triggers: Defending Against Textual Backdoor Attacks with Benign Backdoor Augmentation
- Test-Time Detection of Backdoor Triggers for Poisoned Deep Neural Networks
- Kallima: A Clean-label Framework for Textual Backdoor Attacks
- Clean-Label Backdoor Attacks on Video Recognition Models
- BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input Detection
- On the Vulnerability of Backdoor Defenses for Federated Learning
- Backdoor Attack in the Physical World
- Backdoor Mitigation by Correcting the Distribution of Neural Activations
- Physical Adversarial Attacks on Deep Neural Networks for Traffic Sign Recognition: A Feasibility Study