MAttNet: Modular Attention Network for Referring Expression Comprehension
Explore this paper's citation graph
- Type
- preprint
- Published
- 2018-01-24
- Cited by
- 984
- References
- 36
- Access
- Open access
- OpenAlex
- https://openalex.org/W2784458614
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3441497
Keywords
Computer science, Modular design, Comprehension, Focus (optics), Phrase
References
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fully convolutional networks for semantic segmentation
- Parsing with Compositional Vector Grammars
- Generation and Comprehension of Unambiguous Object Descriptions
- Deep Residual Learning for Image Recognition
- ReferItGame: Referring to Objects in Photographs of Natural Scenes
- Image Captioning with Semantic Attention
- Neural Module Networks
- Hierarchical Attention Networks for Document Classification
- Modeling Context Between Objects for Referring Expression Understanding
- Boosting Image Captioning with Attributes
- Modeling Relationships in Referential Expressions with Compositional Modular Networks
- Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
- A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
- Comprehension-Guided Referring Expressions
- An Implementation of Faster RCNN with Study for Region Sampling
- Recurrent Multimodal Interaction for Referring Image Segmentation
- Inferring and Executing Programs for Visual Reasoning
Cited by
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- Video Object Segmentation with Language Referring Expressions
- Weakly Supervised Attention Learning for Textual Phrases Grounding
- From image to language and back again
- Temporally Grounding Natural Sentence in Video
- Pay Attention! - Robustifying a Deep Visuomotor Policy Through Task-Focused Visual Attention
- Visual Spatial Attention Network for Relationship Detection
- Cross-modal Moment Localization in Videos
- Structural-attentioned LSTM for action recognition based on skeleton
- SEIGAN: Towards Compositional Image Generation by Simultaneously Learning to Segment, Enhance, and Inpaint
- Towards Human-Friendly Referring Expression Generation
- Multi-Level Multimodal Common Semantic Space for Image-Phrase Grounding
- Localizing Natural Language in Videos
- Real-Time Referring Expression Comprehension by Single-Stage Grounding Network
- Explainability by Parsing: Neural Module Tree Networks for Natural Language Visual Grounding
- Neighbourhood Watch: Referring Expression Comprehension via Language-Guided Graph Attention Networks
- Composing Text and Image for Image Retrieval - an Empirical Odyssey
- DeepPhos: prediction of protein phosphorylation sites with deep learning
- CLEVR-Ref+: Diagnosing Visual Reasoning With Referring Expressions
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
Related papers
- Bounding-box Centralization for Improving SiamFC++
- Syncretic-NMS: A Merging Non-Maximum Suppression Algorithm for Instance Segmentation
- Medical image segmentation with imperfect 3D bounding boxes
- A Method to Generate the Minimum Bounding Boxes for Shape-Arbitrary Objects
- Fully Functional Image Manipulation Using Scene Graphs in A Bounding-Box Free Way
- Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint
- Improving Head Pose Estimation with a Combined Loss and Bounding Box Margin Adjustment