FiLM: Visual Reasoning with a General Conditioning Layer
Explore this paper's citation graph
Summary
It is shown that FiLM layers are highly effective for visual reasoning - answering image-related questions which require a multi-step, high-level process - a task which has proven difficult for standard deep learning methods that do not explicitly model reasoning.
- Type
- preprint
- Published
- 2017-09-01
- Cited by
- 4,489
- References
- 44
- Access
- Open access
- OpenAlex
- https://openalex.org/W2734498959
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:19119291
Keywords
Affine transformation, Computer science, Benchmark (surveying), Feature (linguistics), Artificial intelligence
References
- Learning Factored Representations in a Deep Mixture of Experts
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Traversing Knowledge Graphs in Vector Space
- Long Short-Term Memory
- ImageNet Large Scale Visual Recognition Challenge
- Translating Embeddings for Modeling Multi-relational Data
- Ask Your Neurons: A Neural-Based Approach to Answering Questions about Images
- A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
- Distributed Representations of Words and Phrases and their Compositionality
- Visualizing Data using t-SNE
- Deep Residual Learning for Image Recognition
- Neural Module Networks
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- A Learned Representation For Artistic Style
- Overcoming catastrophic forgetting in neural networks
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- VQA: Visual Question Answering
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization
Cited by
- Inferring and Executing Programs for Visual Reasoning
- Deep learning evaluation using deep linguistic processing
- Priming Neural Networks
- Learning by Asking Questions
- Broadcasting Convolutional Network
- DDRprog: A CLEVR Differentiable Dynamic Reasoning Programmer
- Compositional Attention Networks for Machine Reasoning
- Challenging Images For Minds and Machines
- Learning to Count Objects in Natural Images for Visual Question Answering
- A dataset and architecture for visual reasoning with a working memory
- Guide Me: Interacting with Deep Networks
- Learning to Color from Language
- Learning to Follow Language Instructions with Adversarial Reward Induction
- A New Framework for Machine Intelligence: Concepts and Prototype
- TAFE-Net: Task-Aware Feature Embeddings for Efficient Learning and Inference.
- A Question-Answering framework for plots using Deep learning
- ReConvNet: Video Object Segmentation with Spatio-Temporal Features Modulation
- Uncertainty in Multitask Transfer Learning
- Talk the Walk: Navigating New York City through Grounded Dialogue
- Feature-wise transformations
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Solutions to the Third Benchmark Control Problem
- A Benchmark Characterization of the EEMBC Benchmark Suite
- The Performance Validation of Linear Programming Algorithm Based on Integrated Benchmark
- An empirical assessment of Bellon's clone benchmark
- SHREC 2011: robust feature detection and description benchmark
- Gated factored 3-way RBM for image transformation