A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching
Explore this paper's citation graph
- Type
- article
- Published
- 2013-06-23
- Cited by
- 333
- References
- 37
- Access
- Open access
- OpenAlex
- https://openalex.org/W1995820507
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:12284555
Keywords
Computer science, Image stitching, Artificial intelligence, Natural language processing, Natural language
References
- Efficiently Scaling up Crowdsourced Video Annotation
- Topic Models for Image Annotation and Text Illustration
- Comparing Automatic and Human Evaluation of NLG Systems
- Corpus-Guided Sentence Generation of Natural Images
- Topic regression multi-modal Latent Dirichlet Allocation for image annotation
- Modeling annotated data
- A Spatio-Temporal Descriptor Based on 3D-Gradients
- Unbiased look at dataset bias
- The Pascal Visual Object Classes (VOC) Challenge
- Recognition using visual phrases
- Towards coherent natural language description of video streams
- Action bank: A high-level representation of activity in video
- Baby talk: Understanding and generating simple image descriptions
- Translating related words to videos and back through latent topics
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Rethinking LDA: Why Priors Matter
- Collecting Image Annotations Using Amazon’s Mechanical Turk
- Graphical Models, Exponential Families, and Variational Inference
- Spatially Coherent Latent Topic Model for Concurrent Segmentation and Classification of Objects and Scenes
- Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics
Cited by
- Generalized Conditional Matching Algorithm for Ordered and Unordered Sets
- Unsupervised Semantic Parsing of Video Collections
- Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified Framework
- Interpretable video representation
- Combining visual recognition and computational linguistics : linguistic knowledge for visual recognition and natural language descriptions of visual content
- Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide Web
- Generating Multi-Sentence Lingual Descriptions of Indoor Scenes
- Temporally coherent interpretations for long videos using pattern theory
- Book2Movie: Aligning video scenes with book chapters
- Compositional Structure Learning for Action Understanding
- Joint photo stream and blog post summarization and exploration
- Grasp type revisited: A modern perspective on a classical feature for vision
- Long-term recurrent convolutional networks for visual recognition and description
- Deep correlation for matching images and text
- Adopting Abstract Images for Semantic Scene Understanding
- A framework for creating natural language descriptions of video streams
- DISCOVER: Discovering Important Segments for Classification of Video Events and Recounting
- Zero-Shot Event Detection Using Multi-modal Fusion of Weakly Supervised Concepts
- Multimedia Topic Models Considering Burstiness of Local Features
- A maximal figure-of-merit learning approach to maximizing mean average precision with deep neural network based classifiers
Related papers
- Evaluation in Natural Language Generation: Lessons from Referring Expression Generation
- Evaluation in the context of natural language generation
- Affective Natural Language Generation
- Context-aware Natural Language Generation for Spoken Dialogue Systems
- Data mining via protoform based linguistic summaries: Some possible relations to natural language generation
- Integrating natural language generation and model-based reasoning for explanation generation
- Intelligent Help Facilities: Generating Natural Language Descriptions with Examples
- Few-Shot Natural Language Generation by Rewriting Templates