Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences
Explore this paper's citation graph
Summary
This work introduces a multi-level aligner that empowers the alignment-based encoder-decoder model with long short-term memory recurrent neural networks (LSTM-RNN) to translate natural language instructions to action sequences based upon a representation of the observable world state.
- Type
- article
- Published
- 2015-06-12
- Cited by
- 247
- References
- 36
- Access
- Open access
- OpenAlex
- https://openalex.org/W1933065844
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:979457
Keywords
Computer science, Sentence, Recurrent neural network, Benchmark (surveying), Task (project management)
References
- Learning models for following natural language directions in unknown environments
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Procedures As A Representation For Data In A Computer Program For Understanding Natural Language
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Recurrent Neural Network Regularization
- Walk the Talk: Connecting Language, Knowledge, and Action in Route Instructions
- A Neural Attention Model for Abstractive Sentence Summarization
- Show and tell: A neural image caption generator
- Mind's eye: A recurrent visual representation for image caption generation
- Long-term recurrent convolutional networks for visual recognition and description
- Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
- Alignment-Based Compositional Semantics for Instruction Following
- What Are You Talking About? Text-to-Image Coreference
- Long Short-Term Memory
- Following directions using statistical machine translation
- Dropout: a simple way to prevent neural networks from overfitting
- The Symbol Grounding Problem
- Learning to Interpret Natural Language Navigation Instructions from Observations
- Unsupervised PCFG Induction for Grounded Language Learning with Highly Ambiguous Supervision
- Learning to Connect Language and Perception
Cited by
- Articulated Motion Learning via Visual and Lingual Signals
- Survey on the attention based RNN model and its applications in computer vision
- Data Recombination for Neural Semantic Parsing
- Learning Articulated Motion Models from Visual and Lingual Signals
- Semantic Parsing with Semi-Supervised Sequential Autoencoders
- Navigational Instruction Generation as Inverse Reinforcement Learning with Neural Machine Translation
- Coherent Dialogue with Attention-Based Language Models
- Learning to Interpret and Generate Instructional Recipes
- Continuously Improving Natural Language Understanding for Robotic Systems through Semantic Parsing, Dialog, and Multi-modal Perception
- A review of spatial reasoning and interaction for real-world robotics
- Real-time natural language corrections for assistive robotic manipulators
- Multimodal Machine Learning: A Survey and Taxonomy
- Gated-Attention Architectures for Task-Oriented Language Grounding
- Source-Target Inference Models for Spatial Instruction Understanding
- Recurrent neural networks: methods and applications to non-linear predictions
- Domain Transfer for Deep Natural Language Generation from Abstract Meaning Representations
- May I take your order? A Neural Model for Extracting Structured Information from Conversations
- Learning to Generate Market Comments from Stock Prices
- A Tale of Two DRAGGNs: A Hybrid Approach for Interpreting Action-Oriented and Goal-Oriented Instructions
- Land Cover Classification from Multi-temporal, Multi-spectral Remotely Sensed Imagery using Patch-Based Recurrent Neural Networks
Related papers
- Keeping Conflicts Latent: "Salient" versus "Non-Salient" Interpersonal Conflict Management Strategies of Japanese
- A Study on Performance Improvement of Recurrent Neural Networks Algorithm using Word Group Expansion Technique
- Fusion Recurrent Neural Network
- A Computational Method to Find Salient Features
- Gated Feedback Recurrent Neural Networks
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network