Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
Explore this paper's citation graph
Summary
DeepLIFT (Learning Important FeaTures), an efficient and effective method for computing importance scores in a neural network that compares the activation of each neuron to its 'reference activation' and assigns contribution scores according to the difference.
- Type
- preprint
- Published
- 2016-05-05
- Cited by
- 889
- References
- 7
- Access
- Open access
- OpenAlex
- https://openalex.org/W2346578521
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:8564234
Keywords
Interpretability, Black box, Artificial neural network, Computer science, Deep neural networks
References
- On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Long Short-Term Memory
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Deconvolutional networks
- Striving for Simplicity: The All Convolutional Net
- Visualizing and Understanding Convolutional Networks
Cited by
- deepTarget
- An unexpected unity among methods for interpreting model predictions
- Gradients of Counterfactuals
- AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech
- Interpreting the Predictions of Complex ML Models by Layer-wise Relevance Propagation
- Impact of regulatory variation across human iPSCs and differentiated cells
- Predicting Transcriptional Regulatory Activities with Deep Convolutional Networks
- Nucleotide sequence and DNaseI sensitivity are predictive of 3D chromatin architecture
- Reverse-complement parameter sharing improves deep learning models for genomics
- Understanding sequence conservation with deep learning
- Maximum entropy methods for extracting the learned features of deep neural networks
- Axiomatic Attribution for Deep Networks
- Right for the Right Reasons: Training Differentiable Models by Constraining their Explanations
- Learning Important Features Through Propagating Activation Differences
- Using Neural Networks to Improve Single Cell RNA-Seq Data Analysis
- Understanding the Feedforward Artificial Neural Network Model From the Perspective of Network Flow
- Interpreting Blackbox Models via Model Extraction
- MAGIX: Model Agnostic Globally Interpretable Explanations
- RIDDLE: Race and ethnicity Imputation from Disease history with Deep LEarning
- Visual Explanations for Convolutional Neural Networks via Input Resampling
Related papers
- Demystifying Black Box Models with Neural Networks for Accuracy and Interpretability of Supervised Learning
- Thermodynamics-inspired explanations of artificial intelligence
- A Theory of Diagnostic Interpretation in Supervised Classification
- Innovative approaches to addressing the tradeoff between interpretability and accuracy in ship fuel consumption prediction
- Critical Empirical Study on Black-box Explanations in AI
- Hybrid Predictive Model: When an Interpretable Model Collaborates with a Black-box Model
- Unwrapping The Black Box of Deep ReLU Networks: Interpretability, Diagnostics, and Simplification
- Interpretable Mesomorphic Networks for Tabular Data