Situation Recognition: Visual Semantic Role Labeling for Image Understanding
Explore this paper's citation graph
- Type
- article
- Published
- 2016-06-01
- Cited by
- 286
- References
- 55
- OpenAlex
- https://openalex.org/W2423576022
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:2424223
Keywords
FrameNet, Lexicon, Computer science, Clipping (morphology), Artificial intelligence
References
- From TreeBank to PropBank
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- TUHOI: Trento Universal Human Object Interaction Dataset
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question
- Visual Semantic Role Labeling
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Show and tell: A neural image caption generator
- From captions to visual concepts and back
- Describing Common Human Visual Actions in Images
- CIDEr: Consensus-based image description evaluation
- A survey on still image based human action recognition
- Recognizing human actions in still images: a study of bag-of-features and part-based representations
- Background to Framenet
- The Pascal Visual Object Classes Challenge: A Retrospective
- Human action recognition by learning bases of action attributes and parts
- Modeling mutual context of object and human pose in human-object interaction activities
- What Are You Talking About? Text-to-Image Coreference
- Grouplet: A structured image representation for recognizing human and object interactions
Cited by
- Subjects and Their Objects: Localizing Interactees for a Person-Centric View of Importance
- Frames in places: visual common sense knowledge in context
- Active Object Localization in Visual Situations
- Learning to generalize to new compositions in image understanding
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- Computer Vision and Natural Language Processing
- Scene Graph Generation by Iterative Message Passing
- Asynchronous Temporal Fields for Action Recognition
- Recurrent Models for Situation Recognition
- An Analysis of Action Recognition Datasets for Language and Vision Tasks
- A Domain Based Approach to Social Relation Recognition
- Captioning Videos Using Large-Scale Image Corpus
- The “Something Something” Video Database for Learning and Evaluating Visual Common Sense
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1, 600 Papers Survey
- Zero-Shot Activity Recognition with Verb Attribute Induction
- Whodunnit? Crime Drama as a Case for Natural Language Understanding
- Semantic Image Retrieval via Active Grounding of Visual Situations
- Neural Motifs: Scene Graph Parsing with Global Context
- Structured Set Matching Networks for One-Shot Part Labeling
- Disambiguating Visual Verbs
Related papers
- Using Ontologies for Semi-automatic Linking VerbaLex with FrameNet
- Shallow Semantic Parsing Based on FrameNet, VerbNet and PropBank
- A Comparative Study on Generalization of Semantic Roles in FrameNet
- FrameNet and Information Processing in Chinese
- Senseval-3 task: Automatic labeling of semantic roles
- Cyberbullying Lexicon for Social Media
- Building a Spanish lexicon for corpus analysis
- Lexical Acquisition for Opinion Inference: A Sense-Level Lexicon of Benefactive and Malefactive Events