KVT: k-NN Attention for Boosting Vision Transformers
Explore this paper's citation graph
Summary
A sparse attention scheme, dubbed k-NN attention, which naturally inherits the local bias of CNNs without introducing convolutional operations, and allows for the exploration of long range correlation and filter out irrelevant tokens by choosing the most similar tokens from the entire image.
- Type
- preprint
- Published
- 2021-05-28
- Cited by
- 153
- References
- 93
- Access
- Open access
- OpenAlex
- https://openalex.org/W3169938586
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:235266033
Keywords
Computer science, Locality, Transformer, Artificial intelligence, Convolutional neural network
References
- Effective Approaches to Attention-based Neural Machine Translation
- ImageNet Large Scale Visual Recognition Challenge
- Discussion of “Sure Independence Screening for Ultra-High Dimensional Feature Space
- Deep Residual Learning for Image Recognition
- Semantic Understanding of Scenes Through the ADE20K Dataset
- Aggregated Residual Transformations for Deep Neural Networks
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Generating Wikipedia by Summarizing Long Sequences
- MnasNet: Platform-Aware Neural Architecture Search for Mobile
- Generating Long Sequences with Sparse Transformers
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Adaptively Sparse Transformers
- Efficient Content-Based Sparse Attention with Routing Transformers
- Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
- Axial Attention in Multidimensional Transformers
- Sparse Sinkhorn Attention
- Longformer: The Long-Document Transformer
- End-to-End Object Detection with Transformers
- Exploring Self-Attention for Image Recognition
Cited by
- Augmented Transformer with Adaptive Graph for Temporal Action Proposal Generation
- A Survey on Vision Transformer
- Scaled ReLU Matters for Training Vision Transformers
- VSA: Learning Varied-Size Window Attention in Vision Transformers
- DynaST: Dynamic Sparse Transformer for Exemplar-Guided Image Generation
- MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-Spoofing
- Sufficient Vision Transformer
- ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers
- Effective Vision Transformer Training: A Data-Centric Perspective
- The Lottery Ticket Hypothesis for Vision Transformers
- Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion Recognition
- Convolutional Neural Networks or Vision Transformers: Who Will Win the Race for Action Recognitions in Visual Data?
- MFCFlow: A Motion Feature Compensated Multi-Frame Recurrent Network for Optical Flow Estimation
- Local Selective Vision Transformer for Depth Estimation Using a Compound Eye Camera
- Power Time Series Forecasting by Pretrained LM
- Revisit Parameter-Efficient Transfer Learning: A Two-Stage Paradigm
- A Unified Multimodal De- and Re-Coupling Framework for RGB-D Motion Recognition
- What Limits the Performance of Local Self-attention?
- Change Detection of High-Resolution Remote Sensing Images Through Adaptive Focal Modulation on Hierarchical Feature Maps
- Alternating attention Transformer for single image deraining
Related papers
- Clutter rejection limitations from ambiguous range clutter
- Review of visual clutter and its effects on pilot performance: A new look at past research
- Evidence of Clutter Avoidance in Complex Scenes
- Method for Estimating Representative Values of Clutter Heights for Recommendation ITU-R P.452
- Stationary clutter rejection in echocardiography.
- Researching on combining boosting ensembles
- A Brief Introduction to Boosting
- Polarisation behaviour of ground clutter during dwell time
- Ground clutter simulation for surface-based radars