MLP-Mixer: An all-MLP Architecture for Vision
Explore this paper's citation graph
Summary
It is shown that while convolutions and attention are both sufficient for good performance, neither of them are necessary, and MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs), attains competitive scores on image classification benchmarks.
- Type
- preprint
- Published
- 2021-05-04
- Cited by
- 3,780
- References
- 65
- Access
- Open access
- OpenAlex
- https://openalex.org/W3157506437
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:233714958
Keywords
Computer science, Artificial intelligence, Convolutional neural network, Perceptron, Pattern recognition (psychology)
References
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Care of aged doctors
- Acceleration of stochastic approximation by averaging
- Dropout: a simple way to prevent neural networks from overfitting
- Going deeper with convolutions
- ImageNet: A large-scale hierarchical image database
- Backpropagation Applied to Handwritten Zip Code Recognition
- ImageNet classification with deep convolutional neural networks
- Rethinking the Inception Architecture for Computer Vision
- Deep Residual Learning for Image Recognition
- Xception: Deep Learning with Depthwise Separable Convolutions
- Automated Flower Classification over a Large Number of Classes
- Aggregated Residual Transformations for Deep Neural Networks
- Understanding the Effective Receptive Field in Deep Convolutional Neural Networks
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Squeeze-and-Excitation Networks
- Exploring the Limits of Weakly Supervised Pretraining
- AutoAugment: Learning Augmentation Policies from Data
- Gaussian Error Linear Units (GELUs)
- Generating Long Sequences with Sparse Transformers
Cited by
- Are All Layers Created Equal?
- Synthesizer: Rethinking Self-Attention for Transformer Models
- Characterizing Structural Regularities of Labeled Data in Overparameterized Models
- Efficient Transformers: A Survey
- Machine Learning for Cataract Classification/Grading on Ophthalmic Imaging Modalities: A Survey
- Scaling Out-of-Distribution Detection for Real-World Settings
- Transformers in Vision: A Survey
- Red Alarm for Pre-trained Models: Universal Vulnerabilities by Neuron-Level Backdoor Attacks
- ConViT: improving vision transformers with soft convolutional inductive biases
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- A Practical Survey on Faster and Lighter Transformers
- FNet: Mixing Tokens with Fourier Transforms
- ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training
- Automated ASD detection using hybrid deep lightweight features extracted from EEG signals
- Can attention enable MLPs to catch up with CNNs?
- Knowledge distillation: A good teacher is patient and consistent
- MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis
- Identification of Plant-Leaf Diseases Using CNN and Transfer-Learning Approach
- Modulating Language Models with Emotions
- Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual Cortex
Related papers
- An Adaptable Real-Time Object Detection for Traffic Surveillance using R-CNN over CNN with Improved Accuracy
- Cooperative coevolution of generalized multi-layer perceptrons
- POSITIVE-NEGATIVE ASYMMETRY IN MENTAL STATE INFERENCE: REPLICATION AND EXTENSION
- Cooperative Coevolution of Generalized Multi-Layer Perceptrons
- Convolutional Neural Network (CNN) Applied to the Risk Analysis of Accidents in Vessels Navigating the Amazon Rivers
- AdapterDrop: On the Efficiency of Adapters in Transformers
- AdapterDrop: On the Efficiency of Adapters in Transformers