Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Explore this paper's citation graph
Summary
A contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning and proposes momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model.
- Type
- preprint
- Published
- 2021-07-16
- Cited by
- 2,925
- References
- 64
- Access
- Open access
- OpenAlex
- https://openalex.org/W3184735396
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:236034189
Keywords
Computer science, Artificial intelligence, Inference, Encoder, Transformer
References
- Distilling the Knowledge in a Neural Network
- A large annotated corpus for learning natural language inference
- Im2Text: Describing Images Using 1 Million Captioned Photographs
- ReferItGame: Referring to Objects in Photographs of Natural Scenes
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- VQA: Visual Question Answering
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
- Deep Mutual Learning
- MAttNet: Modular Attention Network for Referring Expression Comprehension
- Born Again Neural Networks
- Representation Learning with Contrastive Predictive Coding
- Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
- A Corpus for Reasoning about Natural Language Grounded in Photographs
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Modeling Context in Referring Expressions
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Large scale distributed neural network training through online distillation
- Deep Modular Co-Attention Networks for Visual Question Answering
Cited by
- A Primer on Contrastive Pretraining in Language Processing: Methods, Lessons Learned, and Perspectives
- Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval
- Playing Lottery Tickets with Vision and Language
- Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
- Visual Emotion Representation Learning via Emotion-Aware Pre-training
- VLDeformer: Vision-Language Decomposed Transformer for fast cross-modal retrieval
- A Broad Study of Pre-training for Domain Generalization and Adaptation
- A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning
- CSN: Component-Supervised Network for Few-Shot Classification
- LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval
- Single-Stream Multi-Level Alignment for Vision-Language Pretraining
- WuDaoMM: A large-scale Multi-Modal Dataset for Pre-training models
- Image Captioning In the Transformer Age
- K-LITE: Learning Transferable Visual Models with External Knowledge
- ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models
- PreTraM: Self-Supervised Pre-training via Connecting Trajectory and Map
- Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks
- Training and challenging models for text-guided fashion image retrieval
- Keep the Caption Information: Preventing Shortcut Learning in Contrastive Image-Caption Retrieval
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning