Fooling Neural Network Interpretations via Adversarial Model Manipulation

Explore this paper's citation graph

Summary

It is claimed that the stability of neural network interpretation method with respect to the authors' adversarial model manipulation is an important criterion to check for developing robust and reliable neural network interpretations method.

Type
preprint
Published
2019-02-01
Cited by
246
References
54
Access
Open access

Keywords

Interpretation (philosophy), Computer science, Adversarial system, Artificial intelligence, Set (abstract data type)

References

Cited by

Related papers