From Deterministic to Generative: Multimodal Stochastic RNNs for Video Captioning
Explore this paper's citation graph
- Type
- preprint
- Published
- 2017-08-08
- Cited by
- 232
- References
- 76
- Access
- Open access
- OpenAlex
- https://openalex.org/W2742570236
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:27767604
Keywords
Closed captioning, Computer science, Recurrent neural network, Encoder, Generative model
References
- ADADELTA: An Adaptive Learning Rate Method
- Limited Discrepancy Beam Search
- Recurrent neural network based language model
- A Recurrent Latent Variable Model for Sequential Data
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Describing Videos by Exploiting Temporal Structure
- Natural Language Description of Human Activities from Video Images Based on Concept Hierarchy of Actions
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- From captions to visual concepts and back
- CIDEr: Consensus-based image description evaluation
- Auto-Encoding Variational Bayes
- Discovering Discriminative Graphlets for Aerial Image Categories Recognition
- Human Focused Video Description
- Long Short-Term Memory
- An overview of methods to evaluate uncertainty of deterministic models in decision support
- Going deeper with convolutions
- Bleu: a Method for Automatic Evaluation of Machine Translation
Cited by
- Self-Supervised Video Hashing With Hierarchical Binary Auto-Encoder
- COCO-CN for Cross-Lingual Image Tagging, Captioning, and Retrieval
- Video Captioning with Boundary-aware Hierarchical Language Decoding and Joint Video Prediction
- Pseudo Transfer with Marginalized Corrupted Attribute for Zero-shot Learning
- Relation Classification via LSTMs Based on Sequence and Tree Structure
- A Fine-Grained Spatial-Temporal Attention Model for Video Captioning
- Low-rank dimensionality reduction for multi-modality neurodegenerative disease identification
- Structured Two-Stream Attention Network for Video Question Answering
- Two-stage deep learning for supervised cross-modal retrieval
- Reconstructing Perceived Images From Human Brain Activities With Bayesian Deep Multiview Learning
- Natural Language Generation Using Dependency Tree Decoding for Spoken Dialog Systems
- Leveraging unpaired out-of-domain data for image captioning
- A retrieval algorithm of encrypted speech based on short-term cross-correlation and perceptual hashing
- Multi-scale aggregation network for temporal action proposals
- Photographic painting style transfer using convolutional neural networks
- Semantic consistent adversarial cross-modal retrieval exploiting semantic similarity
- IARNN-Based Semantic-Containing Double-Level Embedding Bi-LSTM for Question-and-Answer Matching
- Exploiting weak mask representation with convolutional neural networks for accurate object tracking
- A new steganalysis approach with an efficient feature selection and classification algorithms for identifying the stego images
- Information fusion in visual question answering: A Survey
Related papers
- TC-VAE: Uncovering Out-of-Distribution Data Generative Factors
- Generative Model for Person Re-Identification: A Review
- Towards Understanding the Interplay of Generative Artificial Intelligence and the Internet
- Are generative approaches to ZSAR a look in the right direction?
- A Comprehensive Review of the Latest Advancements in Large Generative AI Models
- Generating Realistic Blood-Cell Images using Cycle-Consistent Generative Adversial Networks
- Dual-Teacher Class-Incremental Learning With Data-Free Generative Replay
- Deep Dexterous Grasping of Novel Objects from a Single View
- Increasing the Diversity of Deep Generative Models