Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Explore this paper's citation graph

Summary

A contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning and proposes momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model.

Type
preprint
Published
2021-07-16
Cited by
2,925
References
64
Access
Open access

Keywords

Computer science, Artificial intelligence, Inference, Encoder, Transformer

References

Cited by

Related papers