Learning to Count Objects in Natural Images for Visual Question Answering
Explore this paper's citation graph
Summary
A neural network component is proposed that allows robust counting from object proposals and is obtained state-of-the-art accuracy on the number category of the VQA v2 dataset without negatively affecting other categories, even outperforming ensemble models with the authors' single model.
- Type
- preprint
- Published
- 2018-02-15
- Cited by
- 225
- References
- 34
- Access
- Open access
- OpenAlex
- https://openalex.org/W2787119853
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3358859
Keywords
Question answering, Natural (archaeology), Computer science, Artificial intelligence, Closed-ended question
References
- Long Short-Term Memory
- Dropout: a simple way to prevent neural networks from overfitting
- Learning to combine foveal glimpses with a third-order Boltzmann machine
- Learning To Count Objects in Images
- On the Properties of Neural Machine Translation: Encoder–Decoder Approaches
- Counting Everyday Objects in Everyday Scenes
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- VQA: Visual Question Answering
- Structured Attention Networks
- Count-ception: Counting by Fully Convolutional Redundant Counting
- Learning Detection with Diverse Proposals
- Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering
- Learning Non-maximum Suppression
- A simple neural network module for relational reasoning
- Bottom-Up and Top-Down Attention for Image Captioning and VQA
- Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering
Cited by
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- Progressive Reasoning by Module Composition
- Interpretable Visual Question Answering by Reasoning on Dependency Trees
- Chain of Reasoning for Visual Question Answering
- Object-Difference Attention: A Simple Relational Attention for Visual Question Answering
- Fast Parameter Adaptation for Few-shot Image Captioning and Visual Question Answering
- Learning to Compose Dynamic Tree Structures for Visual Contexts
- From Known to the Unknown: Transferring Knowledge to Answer Questions about Novel Visual and Semantic Concepts
- Multi-Task Learning of Hierarchical Vision-Language Representation
- Explainable and Explicit Visual Reasoning Over Scene Graphs
- Learning Representations of Sets through Optimized Permutations
- BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection
- Differential Networks for Visual Question Answering
- The Meaning of “Most” for Visual Question Answering Models
- MUREL: Multimodal Relational Reasoning for Visual Question Answering
- Relation-Aware Graph Attention Network for Visual Question Answering
- What Object Should I Use? - Task Driven Object Detection
- SIMCO: SIMilarity-based object COunting
- Towards VQA Models That Can Read
- Sequential Visual Reasoning for Visual Question Answering
Related papers
- BanglaLM: Data Mining based Bangla Corpus for Language Model Research
- TO ANSWER, OR NOT TO ANSWER, THAT IS THE QUESTION: EXPERTS’ CONTRIBUTION TO QUESTION-ANSWERING PLATFORMS
- Look and Answer the Question: On the Role of Vision in Embodied Question Answering
- Advances in Open-Domain Question Answering
- A visual question answering method based on question intention
- Turkish Question Answering - Question Answering for Distance Education Students
- Natural language question answering: the view from here
- A Weighted Question Retrieval Model using Descriptive Information in Community Question Answering
- Comparing BERT with an intent based question answering setup for open-ended questions in the museum domain