Long-term recurrent convolutional networks for visual recognition and description
Explore this paper's citation graph
- Type
- article
- Published
- 2014-11-17
- Cited by
- 6,469
- References
- 90
- OpenAlex
- https://openalex.org/W1947481528
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5736847
Keywords
Computer science, Artificial intelligence, Benchmark (surveying), Deep learning, Convolutional neural network
References
- Midge: Generating Image Descriptions From Computer Vision Detections
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Generating Text with Recurrent Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Learning to Execute
- Describing Videos by Exploiting Temporal Structure
- Recurrent Neural Network Regularization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Exploring Nearest Neighbor Approaches for Image Captioning
- Generating Sequences With Recurrent Neural Networks
- Corpus-Guided Sentence Generation of Natural Images
- Show and tell: A neural image caption generator
- Beyond short snippets: Deep networks for video classification
- From captions to visual concepts and back
- CIDEr: Consensus-based image description evaluation
- BabyTalk: Understanding and Generating Simple Image Descriptions
- A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching
- Large-Scale Video Classification with Convolutional Neural Networks
Cited by
- Spatiotemporal Recurrent Convolutional Networks for Traffic Prediction in Transportation Networks
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Interactive Sleep Stage Labelling Tool For Diagnosing Sleep Disorder Using Deep Learning
- Employing automatic content recognition for teaching methodology analysis in classroom videos
- Intelligent Data Engineering and Automated Learning – IDEAL 2020: 21st International Conference, Guimaraes, Portugal, November 4–6, 2020, Proceedings, Part II
- Weakly-Supervised Alignment of Video with Text
- Apprentissage de représentations et robotique développementale : quelques apports de l'apprentissage profond pour la robotique autonome. (Representation learning and developmental robotics : on the use of deep learning for autonomous robots)
- How Well Can a CNN Marginalize Simple Nuisances It is Designed for?
- Image Captioning with an Intermediate Attributes Layer
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video
- Differential Recurrent Neural Networks for Action Recognition
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question
- Video Description Generation Incorporating Spatio-Temporal Features and a Soft-Attention Mechanism
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Hard to Cheat: A Turing Test based on Answering Questions about Images
- Visual Semantic Role Labeling
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Solutions to the Third Benchmark Control Problem
- Deep Convolution Neural Network for RBC Images
- Leaf Features Extraction for Plant Classification using CNN
- Small aircraft detection using deep learning
- An Improved Approach for Fire Detection using Deep Learning Models
- Skin Disease Diagnostic techniques using deep learning