A View on Vulnerabilites: The Security Challenges of XAI (Academic Track)
Explore this paper's citation graph
Summary
This survey provides a holistic perspective on the security and safety landscape surrounding XAI, categorizing research on adversarial attacks against XAI and the misuse of explainability to enhance attacks on AI systems, such as evasion and privacy breaches.
- Type
- preprint
- Published
- 2024-01-01
- Cited by
- 5
- References
- 109
- Access
- Open access
- OpenAlex
- https://openalex.org/W2798966449
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:275920423
Keywords
Adversarial system, Computer science, Artificial intelligence, Robustness (evolution), Natural language
References
- On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural Networks
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier
- Towards Evaluating the Robustness of Neural Networks
- SoK: Security and Privacy in Machine Learning
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
- Towards Deep Learning Models Resistant to Adversarial Attacks
- Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR
- One Pixel Attack for Fooling Deep Neural Networks
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- A Survey of Methods for Explaining Black Box Models
- Adversarial Patch
- Model Reconstruction from Model Explanations
- Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)
- Transparency and Explanation in Deep Reinforcement Learning Neural Networks
- Explanations can be manipulated and geometry is to blame
- Generating Natural Language Adversarial Examples
- Interpretation of Neural Networks is Fragile
- A Unified Approach to Interpreting Model Predictions
Cited by
- Early autism detection: a review of emerging technologies, biomarkers, and explainable AI approaches
- AC-TRP: Architecture-Compatible, Causal-Informed Temporal Relevance Propagation Framework for Robust Time Series Analysis
- Explainable But Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
- REaaS: Enabling Adversarially Robust Downstream Classifiers via Robust Encoder as a Service
- A Holistic Approach to Undesired Content Detection in the Real World
- Context-Free Word Importance Scores for Attacking Neural Networks
- Semantic-Preserving Adversarial Text Attacks
- Hierarchical Text Classification with Multi-Label Contrastive Learning and Knn
- SimMC: Simple Masked Contrastive Learning of Skeleton Representations for Unsupervised Person Re-Identification
- CodeAttack: Code-Based Adversarial Attacks for Pre-trained Programming Language Models
- A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
- From Deep Learning to Rational Machines
- A Survey on Image Data Perturbations: Challenges, Techniques, and Impacts
- Attacking Semantic Similarity: Generating Second-Order NLP Adversarial Examples
- Adversarial Training for Improving Model Robustness? Look at Both Prediction and Interpretation
Related papers
- Global Adversarial Attacks for Assessing Deep Learning Robustness
- Developing and Defeating Adversarial Examples
- Generating adversarial examples for DNN using pooling layers
- Adversarial Perturbation Defense on Deep Neural Networks
- A Brief Comparison Between White Box, Targeted Adversarial Attacks in Deep Neural Networks
- Efficient Defenses Against Adversarial Attacks