Fooling Neural Network Interpretations via Adversarial Model Manipulation
Explore this paper's citation graph
Summary
It is claimed that the stability of neural network interpretation method with respect to the authors' adversarial model manipulation is an important criterion to check for developing robust and reliable neural network interpretations method.
- Type
- preprint
- Published
- 2019-02-01
- Cited by
- 246
- References
- 54
- Access
- Open access
- OpenAlex
- https://openalex.org/W2913039310
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:59606176
Keywords
Interpretation (philosophy), Computer science, Adversarial system, Artificial intelligence, Set (abstract data type)
References
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation
- The Emergence of Machine Learning Techniques in Criminology
- Fairness through awareness
- Spearman’s rank correlation coefficient
- ImageNet Large Scale Visual Recognition Challenge
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Deep Residual Learning for Image Recognition
- Evaluating the Visualization of What a Deep Neural Network Has Learned
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier
- Adversarial examples in the physical world
- European Union Regulations on Algorithmic Decision-Making and a "Right to Explanation"
- Deep learning for computational biology
- Grad-CAM: Why did you say that? Visual Explanations from Deep Networks via Gradient-based Localization
- Investigating the influence of noise and distractors on the interpretation of neural networks
- A survey on deep learning in medical image analysis
- Towards A Rigorous Science of Interpretable Machine Learning
- Axiomatic Attribution for Deep Networks
- Learning Important Features Through Propagating Activation Differences
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
Cited by
- A View on Vulnerabilites: The Security Challenges of XAI (Academic Track)
- Towards Hiding Adversarial Examples from Network Interpretation
- Aggregating explainability methods for neural networks stabilizes explanations
- Explanations can be manipulated and geometry is to blame
- Learning Fair Rule Lists
- Fooling Network Interpretation in Image Classification
- Explaining Deep Learning-Based Networked Systems
- How can we fool LIME and SHAP? Adversarial Attacks on Post hoc Explanation Methods
- Benchmarking Attribution Methods with Relative Feature Importance
- Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods
- Aggregating explanation methods for stable and robust explainability.
- Toward Interpretable Machine Learning: Transparent Deep Neural Networks and Beyond
- You Shouldn't Trust Me: Learning Models Which Conceal Unfairness From Multiple Explanation Methods
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
- Adversarial Machine Learning: An Interpretation Perspective
- Gradient Alignment in Deep Neural Networks
- Smoothed Geometry for Robust Attribution
- Fairwashing Explanations with Off-Manifold Detergent
- Adversarial Infidelity Learning for Model Interpretation
- Interpreting Deep Learning-Based Networking Systems