Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Explore this paper's citation graph

Summary

An attention based model that automatically learns to describe the content of images is introduced that can be trained in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound.

Type
article
Published
2015-02-10
Cited by
10,911
References
55
Access
Open access

Keywords

Computer science, Benchmark (surveying), Artificial intelligence, Visualization, Gaze

References

Cited by

Related papers