Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
Explore this paper's citation graph
- Type
- preprint
- Published
- 2017-11-20
- Cited by
- 1,917
- References
- 65
- Access
- Open access
- OpenAlex
- https://openalex.org/W2770387316
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:4673790
Keywords
Computer science, Human–computer interaction, Natural language, Robot, Embodied cognition
References
- Procedures As A Representation For Data In A Computer Program For Understanding Natural Language
- Walk the Talk: Connecting Language, Knowledge, and Action in Route Instructions
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Effective Approaches to Attention-based Neural Machine Translation
- SUN RGB-D: A RGB-D scene understanding benchmark suite
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
- Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences
- Grounding spatial relations for human-robot interaction
- Unbiased look at dataset bias
- Long Short-Term Memory
- Natural language command of an autonomous micro-air vehicle
- Learning to Follow Navigational Directions
- ImageNet Large Scale Visual Recognition Challenge
- Learning to Interpret Natural Language Navigation Instructions from Observations
- Generation and Comprehension of Unambiguous Object Descriptions
- Caffe: Convolutional Architecture for Fast Feature Embedding
- MovieQA: Understanding Stories in Movies through Question-Answering
- Deep Residual Learning for Image Recognition
- Understanding Natural Language Commands for Robotic Navigation and Mobile Manipulation
- ReferItGame: Referring to Objects in Photographs of Natural Scenes
Cited by
- Cognitive Mapping and Planning for Visual Navigation
- Context-aware robot navigation using interactively built semantic maps
- MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments
- Embodied Question Answering
- CHALET: Cornell House Agent Learning Environment
- Look Before You Leap: Bridging Model-Free and Model-Based Reinforcement Learning for Planned-Ahead Vision-and-Language Navigation
- Learning to Navigate in Cities Without a Map
- Situated Mapping of Sequential Instructions to Actions with Single-step Reward Observation
- FollowNet: Robot Navigation by Following Natural Language Directions with Deep Reinforcement Learning
- Points, Paths, and Playscapes: Large-scale Spatial Language Understanding Tasks Set in the Real World
- Following High-level Navigation Instructions on a Simulated Quadcopter with Imitation Learning
- Scheduled Policy Optimization for Natural Language Communication with Intelligent Agents
- Talk the Walk: Navigating New York City through Grounded Dialogue
- Paired Recurrent Autoencoders for Bidirectional Translation Between Robot Actions and Linguistic Descriptions
- Connecting Language and Vision to Actions
- On Evaluation of Embodied Navigation Agents
- Game-Based Video-Context Dialogue
- Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
- Active Vision Dataset Benchmark
- Egocentric Vision-based Future Vehicle Localization for Intelligent Driving Assistance Systems
Related papers
- Deep Residual Learning for Image Recognition
- Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation
- Computer vision and natural language processing for people with vision impairment
- Using vision, acoustics, and natural language for disambiguation
- Reasonable Perception: Connecting Vision and Language Systems for Validating Scene Descriptions
- Work-in-Progress–—Improve Spatial Learning by Chunking Navigation Instructions in Mixed Reality