Modeling Context in Referring Expressions
Explore this paper's citation graph
Summary
This work focuses on incorporating better measures of visual context into referring expression models and finds that visual comparison to other objects within an image helps improve performance significantly.
- Type
- preprint
- Published
- 2016-07-31
- Cited by
- 1,815
- References
- 41
- Access
- Open access
- OpenAlex
- https://openalex.org/W2949107813
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1688357
Keywords
Computer science, Expression (computer science), Focus (optics), Context (archaeology), Natural language processing
References
- Natural Reference to Objects in a Visual Domain
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- LSTM: A Search Space Odyssey
- Show and tell: A neural image caption generator
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- BabyTalk: Understanding and Generating Simple Image Descriptions
- Watching the eyes when talking about size: An investigation of message formulation and utterance planning
- Understanding natural language
- The Use of Spatial Relations in Referring Expression Generation
- Learning Attribute Selections for Non-Pronominal Expressions
- Scalable Object Detection Using Deep Neural Networks
- Im2Text: Describing Images Using 1 Million Captioned Photographs
- ImageNet Large Scale Visual Recognition Challenge
- Computational Generation of Referring Expressions: A Survey
- Generation and Comprehension of Unambiguous Object Descriptions
Cited by
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- Ask Your Neurons: A Deep Learning Approach to Visual Question Answering
- Modeling Context Between Objects for Referring Expression Understanding
- Dense Captioning with Joint Inference and Visual Context
- Modeling Relationships in Referential Expressions with Compositional Modular Networks
- GuessWhat?! Visual Object Discovery through Multi-modal Dialogue
- ImageNet pre-trained models with batch normalization
- A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
- Comprehension-Guided Referring Expressions
- Unsupervised Visual-Linguistic Reference Resolution in Instructional Videos
- Is this a Child, a Girl or a Car? Exploring the Contribution of Distributional Similarity to Learning Referential Word Meanings
- Recurrent Multimodal Interaction for Referring Image Segmentation
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Generating Descriptions with Grounded and Co-referenced People
- Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries
- Deep Reinforcement Learning-Based Image Captioning with Embedding Reward
- Spatio-Temporal Person Retrieval via Natural Language Queries
- Multimodal Machine Learning: A Survey and Taxonomy
- MSRC: multimodal spatial regression with semantic context for phrase grounding
Related papers
- Poetic Experience and Poetic Expression ——On the characteristic of Intuitive Comprehension
- Discussion on the Improvement of Comprehensive and Expression Ability of English Translation
- Evaluation in the context of natural language generation
- Evaluation in Natural Language Generation: Lessons from Referring Expression Generation
- Affective Natural Language Generation
- Context-aware Natural Language Generation for Spoken Dialogue Systems
- Natural Language Generation with Vocabulary Constraints
- Data mining via protoform based linguistic summaries: Some possible relations to natural language generation
- Integrating natural language generation and model-based reasoning for explanation generation
- Intelligent Help Facilities: Generating Natural Language Descriptions with Examples