Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients

Explore this paper's citation graph

Summary

It is demonstrated that regularizing input gradients makes them more naturally interpretable as rationales for model predictions, and also exhibits robustness to transferred adversarial examples generated to fool all of the other models.

Type
article
Published
2017-11-26
Cited by
764
References
33
Access
Open access

Keywords

Interpretability, Adversarial system, Robustness (evolution), Deep neural networks, Artificial intelligence

References

Cited by

Related papers