CoCa: Contrastive Captioners are Image-Text Foundation Models

Explore this paper's citation graph

Summary

Contrastive Captioner (CoCa), a minimalist design to pretrain an image-text encoder-decoder foundation model jointly with contrastive loss and captioning loss, thereby subsuming model capabilities from contrastive approaches like CLIP and generative methods like SimVLM.

Type
preprint
Published
2022-05-04
Cited by
1,815
References
88
Access
Open access

Keywords

Computer science, Closed captioning, Artificial intelligence, Encoder, Deep learning

References

Cited by

Related papers