Inferring the Why in Images
Explore this paper's citation graph
Summary
The results suggest that transferring knowledge from language into vision can help machines understand why a person might be performing an action in an image, and recently developed natural language models to mine knowledge stored in massive amounts of text.
- Type
- preprint
- Published
- 2014-06-20
- Cited by
- 43
- References
- 41
- OpenAlex
- https://openalex.org/W46519926
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:16732164
Keywords
Computer science, Artificial intelligence, Computer vision, Computer graphics (images)
References
- Anticipating Human Activities Using Object Affordances for Reactive Robotic Response
- Task assignment with unknown duration
- NEIL: Extracting Visual Knowledge from Web Data
- BabyTalk: Understanding and Generating Simple Image Descriptions
- Language for learning complex human-object interactions
- A PROCEDURE FOR COMPUTING THE K BEST SOLUTIONS TO DISCRETE OPTIMIZATION PROBLEMS AND ITS APPLICATION TO THE SHORTEST PATH PROBLEM
- Bringing Semantics into Focus Using Visual Abstraction
- Max-Margin Early Event Detectors
- Learning intentions for improved human motion prediction
- Exploiting language models to recognize unseen actions
- Cutting-plane training of structural SVMs
- Human Intent Prediction Using Markov Decision Processes
- The Pascal Visual Object Classes (VOC) Challenge
- Deep networks for predicting human intent with respect to objects
- Predicting human intention in visual observations of hand/object interactions
- Patch to the Future: Unsupervised Visual Prediction
- Visual Persuasion: Inferring Communicative Intents of Images
- Learning Models for Object Recognition from Natural Language Descriptions
- WordNet: A Lexical Database for English
- Learning Everything about Anything: Webly-Supervised Visual Concept Learning
Cited by
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- Don't just listen, use your imagination: Leveraging visual common sense for non-visual tasks
- Grasp type revisited: A modern perspective on a classical feature for vision
- Visual Commonsense for Scene Understanding Using Perception, Semantic Parsing and Reasoning
- MovieQA: Understanding Stories in Movies through Question-Answering
- Learning Common Sense through Visual Abstraction
- Subjects and Their Objects: Localizing Interactees for a Person-Centric View of Importance
- Social LSTM: Human Trajectory Prediction in Crowded Spaces
- Annotation Methodologies for Vision and Language Dataset Creation
- Prediction of Manipulation Actions
- A Corpus for Event Localization
- Learning human activities and poses with interconnected data sources
- Automatic Generation of Grounded Visual Questions
- Asynchronous Temporal Fields for Action Recognition
- Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks
- From Line Drawings to Human Actions: Deep Neural Networks for Visual Data Representation
- From Recognition to Cognition: Visual Commonsense Reasoning
- Leveraging Visual Question Answering for Image-Caption Ranking
- Visual7W: Grounded Question Answering in Images
- Learning the Semantics of Manipulation Action
Related papers
- 6-DOF object localization by combining monocular vision and robot arm kinematics
- 3-D tracking of a moving object by an active stereo vision system
- Self-monitoring to improve robustness of 3D object tracking for robotics
- Object-oriented stripe structured-light vision-guided robot
- Robust object tracking based on RGB-D camera
- An Object Detection and Pose Estimation Approach for Position Based Visual Servoing
- Tracking in 3D: Image Variability Decomposition for Recovering Object Pose and Illumination
- Hand-eye calibration using a single image and robotic picking up using images lacking in contrast
- Motion-based Object Detection and Tracking in Color Image Sequence
- Foreground object segmentation from binocular stereo video