Aligning where to see and what to tell: image caption with region-based attention and scene factorization

Explore this paper's citation graph

Summary

This paper proposes an image caption system that exploits the parallel structures between images and sentences and makes another novel modeling contribution by introducing scene-specific contexts that capture higher-level semantic information encoded in an image.

Type
preprint
Published
2015-06-20
Cited by
121
References
30
Access
Open access

Keywords

Computer science, Artificial intelligence, Salient, Perception, Natural language processing

References

Cited by

Related papers