Generation and Comprehension of Unambiguous Object Descriptions
Explore this paper's citation graph
- Type
- preprint
- Published
- 2015-11-07
- Cited by
- 1,786
- References
- 58
- Access
- Open access
- OpenAlex
- https://openalex.org/W2144960104
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:8745888
Keywords
Closed captioning, Toolbox, Computer science, Object (grammar), Artificial intelligence
References
- Natural Reference to Objects in a Visual Domain
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- A Game-Theoretic Approach to Generating Spatial Descriptions
- Information sharing : reference and presupposition in language generation and interpretation
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Composing Simple Image Descriptions using Web-scale N-grams
- LSTM: A Search Space Odyssey
- Exploring Nearest Neighbor Approaches for Image Captioning
- Corpus-Guided Sentence Generation of Natural Images
- Maximum mutual information estimation of hidden Markov model parameters for speech recognition
- Show and tell: A neural image caption generator
- Mind's eye: A recurrent visual representation for image caption generation
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- CIDEr: Consensus-based image description evaluation
- BabyTalk: Understanding and Generating Simple Image Descriptions
- Understanding natural language
Cited by
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- Resolving References to Objects in Photographs using the Words-As-Classifiers Model
- Reasoning about Pragmatics with Neural Listeners and Speakers
- Ask Your Neurons: A Deep Learning Approach to Visual Question Answering
- Grounding, Justification, Adaptation: Towards Machines That Mean What They Say
- Modeling Context Between Objects for Referring Expression Understanding
- Easy Things First: Installments Improve Referring Expression Generation for Objects in Photographs
- Utilizing Large Scale Vision and Text Datasets for Image Segmentation from Referring Expressions
- Dense Captioning with Joint Inference and Visual Context
- Phrase Localization and Visual Relationship Detection with Comprehensive Linguistic Cues
- Modeling Relationships in Referential Expressions with Compositional Modular Networks
- GuessWhat?! Visual Object Discovery through Multi-modal Dialogue
- ImageNet pre-trained models with batch normalization
- Top-Down Visual Saliency Guided by Captions
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
- Context-Aware Captions from Context-Agnostic Supervision
- Comprehension-Guided Referring Expressions
- Unsupervised Visual-Linguistic Reference Resolution in Instructional Videos
- Statistical models of learning and using semantic representations
Related papers
- OSCAR and ActivityNet: an Image Captioning model can effectively learn a Video Captioning dataset
- Video Captioning via Hierarchical Reinforcement Learning
- Image Captioning Methodologies Using Deep Learning: A Review
- Image Captioning using Neural Networks
- Boosted Attention: Leveraging Human Attention for Image Captioning
- Evolution of Image Captioning Models: An Overview
- Image Captioning- Bangladesh’s Heritage Perspective Using Deep Learning
- Poetic Experience and Poetic Expression ——On the characteristic of Intuitive Comprehension
- Discussion on the Improvement of Comprehensive and Expression Ability of English Translation