Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Explore this paper's citation graph
- Type
- article
- Published
- 2017-07-25
- Cited by
- 4,746
- References
- 66
- Access
- Open access
- OpenAlex
- https://openalex.org/W2745461083
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3753452
Keywords
Closed captioning, Question answering, Computer science, Top-down and bottom-up design, Image (mathematics)
References
- ADADELTA: An Adaptive Learning Rate Method
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- You Only Look Once: Unified, Real-Time Object Detection
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- CIDEr: Consensus-based image description evaluation
- Perceptual grouping and attention in visual search for features and for objects.
- Control of goal-directed and stimulus-driven attention in the brain
- Long Short-Term Memory
- Selective Search for Object Recognition
- Maximum Expected BLEU Training of Phrase and Lexicon Translation Models
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Shifting visual attention between objects and locations: evidence from normal and parietal lesion subjects.
- ImageNet Large Scale Visual Recognition Challenge
- Meteor Universal: Language Specific Translation Evaluation for Any Target Language
- A feature-integration theory of attention.
Cited by
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
- Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering
- Stack-Captioning: Coarse-to-Fine Learning for Image Captioning
- Can you fool AI with adversarial examples on a visual Turing test?
- Convolutional Image Captioning
- Grounded Objects and Interactions for Video Captioning
- Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
- Visual Question Answering as a Meta Learning Task
- Embodied Question Answering
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- Consensus-based Sequence Training for Video Captioning
- Object-Based Reasoning in VQA
- Learning to Count Objects in Natural Images for Visual Question Answering
- Dual Recurrent Attention Units for Visual Question Answering
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- Attention on Attention: Architectures for Visual Question Answering (VQA)
- LSTM stack-based Neural Multi-sequence Alignment TeCHnique (NeuMATCH)
Related papers
- OSCAR and ActivityNet: an Image Captioning model can effectively learn a Video Captioning dataset
- Video Captioning via Hierarchical Reinforcement Learning
- Image Captioning Methodologies Using Deep Learning: A Review
- Image Captioning using Neural Networks
- Boosted Attention: Leveraging Human Attention for Image Captioning
- Image Captioning- Bangladesh’s Heritage Perspective Using Deep Learning
- How Closed Captioning in the U.S. Today can Become the Advanced Television Captioning System of Tomorrow
- Combining bottom-up and top-down attentional influences
- A salient object detection framework beyond top-down and bottom-up mechanism