CoCa: Contrastive Captioners are Image-Text Foundation Models
Explore this paper's citation graph
Summary
Contrastive Captioner (CoCa), a minimalist design to pretrain an image-text encoder-decoder foundation model jointly with contrastive loss and captioning loss, thereby subsuming model capabilities from contrastive approaches like CLIP and generative methods like SimVLM.
- Type
- preprint
- Published
- 2022-05-04
- Cited by
- 1,815
- References
- 88
- Access
- Open access
- OpenAlex
- https://openalex.org/W4229042118
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:248512473
Keywords
Computer science, Closed captioning, Artificial intelligence, Encoder, Deep learning
References
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Neural Machine Translation of Rare Words with Subword Units
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- Fully convolutional networks for semantic segmentation
- A Learning Algorithm for Continually Running Fully Recurrent Neural Networks
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation
- ImageNet: A large-scale hierarchical image database
- Two-Stream Convolutional Networks for Action Recognition in Videos
- ImageNet classification with deep convolutional neural networks
- Deep Residual Learning for Image Recognition
- MSR-VTT: A Large Video Description Dataset for Bridging Video and Language
- Self-Critical Sequence Training for Image Captioning
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- The Kinetics Human Action Video Dataset
- Moments in Time Dataset: One Million Videos for Event Understanding
- Exploring the Limits of Weakly Supervised Pretraining
- A Short Note about Kinetics-600
- Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates
Cited by
- The Computational Limits of Deep Learning
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
- Single-Stream Multi-Level Alignment for Vision-Language Pretraining
- CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks
- Unlocking High-Accuracy Differentially Private Image Classification through Scale
- XnODR and XnIDR: Two Accurate and Fast Fully Connected Layers for Convolutional Neural Networks
- Training Vision-Language Transformers from Captions Alone
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- VL-BEiT: Generative Vision-Language Pretraining
- Neural Collapse: A Review on Modelling Principles and Generalization
- Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training
- Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs
- Visual Clues: Bridging Vision and Language Foundations for Image Paragraph Captioning
- Multimodal Masked Autoencoders Learn Transferable Representations
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts
- Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
- MixGen: A New Multi-Modal Data Augmentation
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- Bridge-Tower: Building Bridges Between Encoders in Vision-Language Representation Learning
- REVECA - Rich Encoder-decoder framework for Video Event CAptioner
Related papers
- ThaiTC:Thai Transformer-based Image Captioning
- Distance Transformer for Image Captioning
- Captioning Transformer with Stacked Attention Modules
- What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning
- EfficientNet-Transformer for image captioning in Bahasa
- Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
- Language Models with Transformers
- Finnish Language Modeling with Deep Transformer Models