BEiT: BERT Pre-Training of Image Transformers

Explore this paper's citation graph

Summary

A self-supervised vision representation model BEiT, which stands for Bidirectional Encoder representation from Image Transformers, is introduced, and results on image classification and semantic segmentation show that the model achieves competitive results with previous pre-training methods.

Type
preprint
Published
2021-06-15
Cited by
3,864
References
59
Access
Open access

Keywords

Computer science, Encoder, Transformer, Artificial intelligence, Pixel

References

Cited by

Related papers