Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Explore this paper's citation graph
Summary
Concept Activation Vectors (CAVs) are introduced, which provide an interpretation of a neural net's internal state in terms of human-friendly concepts, and may be used to explore hypotheses and generate insights for a standard image classification network as well as a medical application.
- Type
- preprint
- Published
- 2017-11-30
- Cited by
- 2,472
- References
- 39
- Access
- Open access
- OpenAlex
- https://openalex.org/W2796885425
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:51737170
Keywords
Interpretability, Computer science, Artificial intelligence, Interpretation (philosophy), Image (mathematics)
References
- Graph-Sparse LDA: A Topic Model with Structured Sparsity
- Intriguing properties of neural networks
- Labeled Faces in the Wild: A Database forStudying Face Recognition in Unconstrained Environments
- Supersparse Linear Integer Models for Interpretable Classification
- Sparse Principal Component Analysis
- Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission
- Machine learning on brain MRI data for differential diagnosis of Parkinson's disease and Progressive Supranuclear Palsy.
- Going deeper with convolutions
- ImageNet Large Scale Visual Recognition Challenge
- Regression Shrinkage and Selection via the Lasso
- Distributed Representations of Words and Phrases and their Compositionality
- Rethinking the Inception Architecture for Computer Vision
- Mind the Gap: A Generative Approach to Interpretable Feature Selection and Extraction
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier
- European Union Regulations on Algorithmic Decision-Making and a "Right to Explanation"
- Grad-CAM: Why did you say that?
- Towards A Rigorous Science of Interpretable Machine Learning
- Axiomatic Attribution for Deep Networks
- Network Dissection: Quantifying Interpretability of Deep Visual Representations
- SmoothGrad: removing noise by adding noise
Cited by
- Embedding Deep Networks into Visual Explanations
- Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
- Productivity, Portability, Performance: Data-Centric Python
- DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems
- Evaluating Feature Importance Estimates
- The challenge of crafting intelligible intelligence
- Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts
- Training Machine Learning Models by Regularizing their Explanations
- Model Cards for Model Reporting
- Secure Deep Learning Engineering: A Software Quality Assurance Perspective
- Scalable agent alignment via reward modeling: a research direction
- A Survey of Evaluation Methods and Measures for Interpretable Machine Learning
- Networks for Nonlinear Diffusion Problems in Imaging
- Can I trust you more? Model-Agnostic Hierarchical Explanations
- Safety and Trustworthiness of Deep Neural Networks: A Survey
- Interactive Naming for Explaining Deep Neural Networks: A Formative Study
- A Learning Effect by Presenting Machine Prediction as a Reference Answer in Self-correction
- Explainable Sentiment Analysis with Applications in Medicine
- Attention-Based Prototypical Learning Towards Interpretable, Confident and Robust Deep Neural Networks
- Explaining Explanations: An Overview of Interpretability of Machine Learning
Related papers
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
- Towards A Rigorous Science of Interpretable Machine Learning
- Sanity Checks for Saliency Maps
- On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation
- Learning Important Features Through Propagating Activation Differences
- Axiomatic Attribution for Deep Networks