Aligning where to see and what to tell: image caption with region-based attention and scene factorization
Explore this paper's citation graph
Summary
This paper proposes an image caption system that exploits the parallel structures between images and sentences and makes another novel modeling contribution by introducing scene-specific contexts that capture higher-level semantic information encoded in an image.
- Type
- preprint
- Published
- 2015-06-20
- Cited by
- 121
- References
- 30
- Access
- Open access
- OpenAlex
- https://openalex.org/W1785460851
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:17326626
Keywords
Computer science, Artificial intelligence, Salient, Perception, Natural language processing
References
- Midge: Generating Image Descriptions From Computer Vision Detections
- Generating Text with Recurrent Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Composing Simple Image Descriptions using Web-scale N-grams
- Corpus-Guided Sentence Generation of Natural Images
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- Long-term recurrent convolutional networks for visual recognition and description
- BabyTalk: Understanding and Generating Simple Image Descriptions
- Long Short-Term Memory
- Selective Search for Object Recognition
- Factored conditional restricted Boltzmann Machines for modeling motion style
- Collecting Image Annotations Using Amazon’s Mechanical Turk
- The Stanford CoreNLP Natural Language Processing Toolkit
- Learning Deep Features for Scene Recognition using Places Database
- Image Description using Visual Dependency Representations
- Collective Generation of Natural Image Descriptions
- Explain Images with Multimodal Recurrent Neural Networks
Cited by
- ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering
- What value high level concepts in vision to language problems
- Survey on the attention based RNN model and its applications in computer vision
- Image Captioning and Visual Question Answering Based on Attributes and Their Related External Knowledge
- What Value Do Explicit High Level Concepts Have in Vision to Language Problems?
- LSTM-in-LSTM for generating long descriptions of images
- Image Captioning with both Object and Scene Information
- Bootstrap, Review, Decode: Using Out-of-Domain Textual Data to Improve Image Captioning
- Dense Captioning with Joint Inference and Visual Context
- Attention-based Memory Selection Recurrent Network for Language Modeling
- Semantic Compositional Networks for Visual Captioning
- Areas of Attention for Image Captioning
- Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
- Diverse Image Captioning via GroupTalk
- Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?
- Reference Based LSTM for Image Captioning
- Image Caption with Global-Local Attention
- AMC: Attention Guided Multi-modal Correlation Learning for Image Search
- Deep Reinforcement Learning-Based Image Captioning with Embedding Reward
- Bottom-Up and Top-Down Attention for Image Captioning and VQA
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Keeping Conflicts Latent: "Salient" versus "Non-Salient" Interpersonal Conflict Management Strategies of Japanese
- Solutions to the Third Benchmark Control Problem
- A Benchmark Characterization of the EEMBC Benchmark Suite
- The Performance Validation of Linear Programming Algorithm Based on Integrated Benchmark
- Word Representation With Salient Features