Variational Structured Semantic Inference for Diverse Image Captioning
Explore this paper's citation graph
Summary
A Variational Structured Semantic Inferring model executed in a novel structured encoder-inferer-decoder schema that achieves significant improvements over the state-of-the-arts in image captioning.
- Type
- article
- Published
- 2019-01-01
- Cited by
- 33
- References
- 35
- OpenAlex
- https://openalex.org/W2970626077
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:202770245
Keywords
Closed captioning, Computer science, Inference, Artificial intelligence, Encoder
References
- Parsing Natural Scenes and Natural Language with Recursive Neural Networks
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Exploring Nearest Neighbor Approaches for Image Captioning
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- Auto-Encoding Variational Bayes
- Approximating the Kullback Leibler Divergence Between Gaussian Mixture Models
- Learning Structured Output Representation using Deep Conditional Generative Models
- Attribute2Image: Conditional Image Generation from Visual Attributes
- Generating Sentences from a Continuous Space
- Image Captioning with Semantic Attention
- Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models
- Context-Aware Captions from Context-Agnostic Supervision
- Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning
- Diverse Image Captioning via GroupTalk
- Towards a Visual Privacy Advisor: Understanding and Predicting Privacy Risks in Images
- Creativity: Generating Diverse Questions Using Variational Autoencoders
- Attend to You: Personalized Image Captioning with Context Sequence Memory Networks
- StyleNet: Generating Attractive Visual Captions with Styles
Cited by
- Information Competing Process for Learning Diversified Representations
- c-TextGen: Conditional Text Generation for Harmonious Human-Machine Interaction
- Boosting image caption generation with feature fusion module
- Indoor Scene Change Captioning Based on Multimodality Data
- Emerging Trends of Multimodal Research in Vision and Language
- Diverse Image Captioning with Context-Object Split Latent Spaces
- Multimodal research in vision and language: A review of current and emerging trends
- Conditional Text Generation for Harmonious Human-Machine Interaction
- Human-like Controllable Image Captioning with Verb-specific Semantic Roles
- Saying the Unseen: Video Descriptions via Dialog Agents
- Diversity as a By-Product: Goal-oriented Language Generation Leads to Linguistic Variation
- Structured Multi-modal Feature Embedding and Alignment for Image-Sentence Retrieval
- Hierarchical Graph Attention Network for Few-shot Visual-Semantic Learning
- Partial Off-policy Learning: Balance Accuracy and Diversity for Human-Oriented Image Captioning
- From Show to Tell: A Survey on Deep Learning-Based Image Captioning
- Learning Distinct and Representative Modes for Image Captioning
- Show, Tell and Rephrase: Diverse Video Captioning via Two-Stage Progressive Training
- Cross-modal Semantic Enhanced Interaction for Image-Sentence Retrieval
- Divcon: Learning Concept Sequences for Semantically Diverse Image Captioning
- Fully-attentive iterative networks for region-based controllable image and video captioning
Related papers
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- ROUGE: A Package for Automatic Evaluation of Summaries
- Bleu: a Method for Automatic Evaluation of Machine Translation
- CIDEr: Consensus-based image description evaluation
- Show and tell: A neural image caption generator
- Deep Residual Learning for Image Recognition
- METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention