Diverse Image Captioning via GroupTalk
Explore this paper's citation graph
Summary
A novel iterative update strategy is proposed to separate training sentence samples into groups and learn their distributions at the same time and introduce an efficient classifier to solve the problem brought about by the non-linear and discontinuous nature of language distributions which will impair performance.
- Type
- article
- Published
- 2016-07-09
- Cited by
- 35
- References
- 25
- OpenAlex
- https://openalex.org/W2578466053
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:6707750
Keywords
Closed captioning, Computer science, Artificial intelligence, Classifier (UML), Image (mathematics)
References
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Image Captioning with an Intermediate Attributes Layer
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- Maximum likelihood from incomplete data via the EM - algorithm plus discussions on the paper
- Deep Compositional Cross-modal Learning to Rank via Local-Global Alignment
- Empirical study of topic modeling in Twitter
- Long Short-Term Memory
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Learning a Recurrent Visual Representation for Image Caption Generation
- Learning Like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Explain Images with Multimodal Recurrent Neural Networks
- Multimodal Neural Language Models
- From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
- A Diversity-Promoting Objective Function for Neural Conversation Models
Cited by
- Towards a Visual Privacy Advisor: Understanding and Predicting Privacy Risks in Images
- From Deterministic to Generative: Multimodal Stochastic RNNs for Video Captioning
- Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning
- GroupCap: Group-Based Image Captioning with Structured Relevance and Diversity Constraints
- A Question Type Driven Framework to Diversify Visual Question Generation
- Measuring the Diversity of Automatic Image Descriptions
- More is Better: Precise and Detailed Image Captioning Using Online Positive Recall and Missing Concepts Mining
- Trainable Decoding of Sets of Sequences for Neural Sequence Models
- Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
- Structured Fusion Networks for Dialog
- Towards Diverse and Accurate Image Captions via Reinforcing Determinantal Point Process
- Variational Structured Semantic Inference for Diverse Image Captioning
- Generating Diverse and Descriptive Image Captions Using Visual Paraphrases
- Latent Normalizing Flows for Many-to-Many Cross-Domain Mappings
- Analysis of diversity-accuracy tradeoff in image captioning
- Multi-Sentence Video Captioning using Content-oriented Beam Searching and Multi-stage Refining Algorithm
- On Diversity in Image Captioning: Metrics and Methods
- Diverse Image Captioning with Context-Object Split Latent Spaces
- Neighbours Matter: Image Captioning with Similar Images
- Diversity as a By-Product: Goal-oriented Language Generation Leads to Linguistic Variation
Related papers
- Bleu: a Method for Automatic Evaluation of Machine Translation
- CIDEr: Consensus-based image description evaluation
- Show and tell: A neural image caption generator
- Meteor Universal: Language Specific Translation Evaluation for Any Target Language
- Deep Residual Learning for Image Recognition
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention