KVT: k-NN Attention for Boosting Vision Transformers

Explore this paper's citation graph

Summary

A sparse attention scheme, dubbed k-NN attention, which naturally inherits the local bias of CNNs without introducing convolutional operations, and allows for the exploration of long range correlation and filter out irrelevant tokens by choosing the most similar tokens from the entire image.

Type
preprint
Published
2021-05-28
Cited by
153
References
93
Access
Open access

Keywords

Computer science, Locality, Transformer, Artificial intelligence, Convolutional neural network

References

Cited by

Related papers