An Empirical Study of Training Self-Supervised Vision Transformers
Explore this paper's citation graph
- Type
- article
- Published
- 2021-04-05
- Cited by
- 2,469
- References
- 50
- Access
- Open access
- OpenAlex
- https://openalex.org/W3145450063
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:233024948
Keywords
Computer science, Artificial intelligence, Machine learning, Transformer, Benchmark (surveying)
References
- One weird trick for parallelizing convolutional neural networks
- Rectified Linear Units Improve Restricted Boltzmann Machines
- Care of aged doctors
- ImageNet: A large-scale hierarchical image database
- Video Google: a text retrieval approach to object matching in videos
- Dimensionality Reduction by Learning an Invariant Mapping
- Backpropagation Applied to Handwritten Zip Code Recognition
- Signature Verification Using A "Siamese" Time Delay Neural Network
- Deep Residual Learning for Image Recognition
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Automated Flower Classification over a Large Number of Classes
- Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Scaling SGD Batch Size to 32K for ImageNet Training
- Large Batch Training of Convolutional Networks
- Unsupervised Feature Learning via Non-parametric Instance Discrimination
- Representation Learning with Contrastive Predictive Coding
- Learning deep representations by mutual information estimation and maximization
- Selective Kernel Networks
- GCNet: Non-Local Networks Meet Squeeze-Excitation Networks and Beyond
Cited by
- Exploring Simple Siamese Representation Learning
- Transformers in Vision: A Survey
- Divide and Contrast: Self-supervised Learning from Uncurated Data
- Salient Objects in Clutter
- Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model
- Self-Supervised Learning of Domain Invariant Features for Depth Estimation
- PASS: An ImageNet replacement for self-supervised pretraining without humans
- Meta-FSEO: A Meta-Learning Fast Adaptation with Self-Supervised Embedding Optimization for Few-Shot Remote Sensing Scene Classification
- A Survey on Vision Transformer
- A Study of the Generalizability of Self-Supervised Representations
- A Low Rank Promoting Prior for Unsupervised Contrastive Learning
- On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals
- VidTr: Video Transformer Without Convolutions
- SSAST: Self-Supervised Audio Spectrogram Transformer
- A Survey of Self-Supervised and Few-Shot Object Detection
- Multi-task vision transformer using low-level chest X-ray feature corpus for COVID-19 diagnosis and severity quantification
- Attention mechanisms in computer vision: A survey
- Are we ready for a new paradigm shift? A survey on visual deep MLP
- LiT: Zero-Shot Transfer with Locked-image text Tuning
- Leveraging Batch Normalization for Vision Transformers
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Solutions to the Third Benchmark Control Problem
- Supervised Machine Learning a Brief Survey of Approaches
- A Supervised Machine Learning Algorithms: Applications, Challenges, and Recommendations
- A Review Paper on A Comparative Study of Supervised Learning Approaches
- Semi-Supervised Learning for Natural Language Processing