Transformers in Vision: A Survey
Explore this paper's citation graph
Summary
This survey aims to provide a comprehensive overview of the Transformer models in the computer vision discipline with an introduction to fundamental concepts behind the success of Transformers, i.e., self-attention, large-scale pre-training, and bidirectional feature encoding.
- Type
- preprint
- Published
- 2021-01-04
- Cited by
- 3,798
- References
- 286
- Access
- Open access
- OpenAlex
- https://openalex.org/W3119997354
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:230435805
Keywords
Computer science, Transformer, Segmentation, Artificial intelligence, Scalability
References
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- You Only Look Once: Unified, Real-Time Object Detection
- Intriguing properties of neural networks
- Distilling the Knowledge in a Neural Network
- Show and tell: A neural image caption generator
- Auto-Encoding Variational Bayes
- Long Short-Term Memory
- A non-local algorithm for image denoising
- Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments
- ImageNet: A large-scale hierarchical image database
- Im2Text: Describing Images Using 1 Million Captioned Photographs
- Generation and Comprehension of Unambiguous Object Descriptions
- Backpropagation Applied to Handwritten Zip Code Recognition
- Semantic object classes in video: A high-definition ground truth database
- Rethinking the Inception Architecture for Computer Vision
- From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
- ShapeNet: An Information-Rich 3D Model Repository
- Deep Residual Learning for Image Recognition
- ReferItGame: Referring to Objects in Photographs of Natural Scenes
Cited by
- ExGate: Externally Controlled Gating for Feature-based Attention in Artificial Neural Networks
- An Attentive Survey of Attention Models
- A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
- Automated Detection and Forecasting of COVID-19 using Deep Learning Techniques: A Review
- Short-Term Traffic Prediction With Deep Neural Networks: A Survey
- Human Action Recognition From Various Data Modalities: A Review
- Curriculum Learning: A Survey
- Classical and Deep Learning based Visual Servoing Systems: a Survey on State of the Art
- Perspectives and Prospects on Transformer Architecture for Cross-Modal Tasks with Language and Vision
- Convolution-Free Medical Image Segmentation using Transformers
- Countering Malicious DeepFakes: Survey, Battleground, and Horizon
- 3D Human Pose Estimation with Spatial and Temporal Transformers
- Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection
- Creativity and Machine Learning: A Survey
- Handwriting Transformers
- On Generating Transferable Targeted Perturbations
- Understanding Robustness of Transformers for Image Classification
- A Practical Survey on Faster and Lighter Transformers
- VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers
- Prototype Memory for Large-scale Face Representation Learning
Related papers
- Deep Residual Learning for Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- ImageNet: A large-scale hierarchical image database
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Training data-efficient image transformers & distillation through attention
- Long Short-Term Memory