VLDeformer: Vision-Language Decomposed Transformer for fast cross-modal retrieval

Explore this paper's citation graph

Summary

A novel Vision-Language Decomposed Transformer (VLDeformer) is presented, which greatly increases the efficiency of VL transformers while maintaining their outstanding accuracy and outperforms state-of-the-art visual-semantic embedding methods on COCO and Flickr30k.

Type
article
Published
2021-10-20
Cited by
26
References
35
Access
Open access

Keywords

Transformer, Computer science, Modal, Search engine indexing, Artificial intelligence

References

Cited by

Related papers