BEiT: BERT Pre-Training of Image Transformers
Explore this paper's citation graph
Summary
A self-supervised vision representation model BEiT, which stands for Bidirectional Encoder representation from Image Transformers, is introduced, and results on image classification and semantic segmentation show that the model achieves competitive results with previous pre-training methods.
- Type
- preprint
- Published
- 2021-06-15
- Cited by
- 3,864
- References
- 59
- Access
- Open access
- OpenAlex
- https://openalex.org/W3170863103
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:235436185
Keywords
Computer science, Encoder, Transformer, Artificial intelligence, Pixel
References
- Neural Machine Translation of Rare Words with Subword Units
- Auto-Encoding Variational Bayes
- ImageNet Large Scale Visual Recognition Challenge
- Semantic Understanding of Scenes Through the ADE20K Dataset
- Unsupervised Representation Learning by Predicting Image Rotations
- Unsupervised Feature Learning via Non-parametric Instance Discrimination
- Representation Learning with Contrastive Predictive Coding
- Learning deep representations by mutual information estimation and maximization
- Cross-lingual Language Model Pretraining
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- Selfie: Self-supervised Pretraining for Image Embedding
- Categorical Reparameterization with Gumbel-Softmax
- Deep Clustering for Unsupervised Learning of Visual Features
- Colorful Image Colorization
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- SpanBERT: Improving Pre-training by Representing and Predicting Spans
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Learning Representations by Maximizing Mutual Information Across Views
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Cited by
- Creativity and Machine Learning: A Survey
- ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training
- Exploring the Diversity and Invariance in Yourself for Visual Pre-Training Task
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- A Survey on Vision Transformer
- Evaluating transformer-based semantic segmentation networks for pathological image segmentation
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
- SSAST: Self-Supervised Audio Spectrogram Transformer
- Attention mechanisms in computer vision: A survey
- Are we ready for a new paradigm shift? A survey on visual deep MLP
- Visio-Linguistic Brain Encoding
- BEVT: BERT Pretraining of Video Transformers
- UNETR: Transformers for 3D Medical Image Segmentation
- Wearable Sensor-Based Human Activity Recognition with Transformer Model
- Fine-Tuning BERT Based Approach for Multi-Class Sentiment Analysis on Twitter Emotion Data
- A Broad Study of Pre-training for Domain Generalization and Adaptation
- Beyond Masking: Demystifying Token-Based Pre-Training for Vision Transformers
- Exploring Plain Vision Transformer Backbones for Object Detection
- In-N-Out Generative Learning for Dense Unsupervised Video Segmentation
- Masked Autoencoders for Point Cloud Self-supervised Learning
Related papers
- Low complexity photo sensor dead pixel detection algorithm
- Sub-pixel mapping based on sub-pixel to sub-pixel spatial attraction model
- Design and Realization of RS Encoder Based on FPGA
- Bad pixel identification by means of principal components analysis
- Sub-pixel mapping of remotely sensed imagery with hybrid intra- and inter-pixel dependence
- Improved K-Pass Pixel Value Ordering Based Data Hiding