Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
Explore this paper's citation graph
Summary
It is demonstrated that regularizing input gradients makes them more naturally interpretable as rationales for model predictions, and also exhibits robustness to transferred adversarial examples generated to fool all of the other models.
- Type
- article
- Published
- 2017-11-26
- Cited by
- 764
- References
- 33
- Access
- Open access
- OpenAlex
- https://openalex.org/W2768346313
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:19167025
Keywords
Interpretability, Adversarial system, Robustness (evolution), Deep neural networks, Artificial intelligence
References
- Intriguing properties of neural networks
- Decolorize: Fast, contrast enhancing, color to grayscale conversion
- Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission
- A Simple Method to Determine if a Music Information Retrieval System is a “Horse”
- Improving generalization performance using double backpropagation
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural Networks
- The Limitations of Deep Learning in Adversarial Settings
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier
- Reading Digits in Natural Images with Unsupervised Feature Learning
- Adversarial examples in the physical world
- Defensive Distillation is Not Robust to Adversarial Examples
- Cleverhans V0.1: an Adversarial Machine Learning Library
- Gradients of Counterfactuals
- Auditing black-box models for indirect influence
- Right for the Right Reasons: Training Differentiable Models by Constraining their Explanations
- Biologically inspired protection of deep networks from adversarial attacks
- Practical Black-Box Attacks against Machine Learning
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
- Prediction of crime occurrence from multi-modal data using deep learning
- The Space of Transferable Adversarial Examples
Cited by
- Adversarial Detection of Flash Malware: Limitations and Open Issues
- Gradient Regularization Improves Accuracy of Discriminative Models
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Adversarial Vulnerability of Neural Networks Increases With Input Dimension
- Understanding and Enhancing the Transferability of Adversarial Examples
- Improving DNN Robustness to Adversarial Attacks using Jacobian Regularization
- Security Consideration For Deep Learning-Based Image Forensics
- Towards Robust Training of Neural Networks by Regularizing Adversarial Gradients
- Evaluating Feature Importance Estimates
- Gradient Band-based Adversarial Training for Generalized Attack Immunity of A3C Path Finding
- Defense Against Adversarial Attacks with Saak Transform
- Adversarial Examples: Attacks on Machine Learning-based Malware Visualization Detection Methods
- On the Intriguing Connections of Regularization, Input Gradients and Transferability of Evasion and Poisoning Attacks
- Adversarial Examples: Opportunities and Challenges
- Why the Failure? How Adversarial Examples Can Provide Insights for Interpretable Machine Learning
- Interpreting Adversarial Robustness: A View from Decision Surface in Input Space
- Stakeholders in Explainable AI
- Training Machine Learning Models by Regularizing their Explanations
- What can AI do for me?: evaluating machine learning interpretations in cooperative play
- Efficient Two-Step Adversarial Defense for Deep Neural Networks
Related papers
- Global Adversarial Attacks for Assessing Deep Learning Robustness
- Developing and Defeating Adversarial Examples
- Adversarial Perturbation Defense on Deep Neural Networks
- Efficient Defenses Against Adversarial Attacks
- Unifying Adversarial Training Algorithms with Flexible Deep Data Gradient Regularization