Learning Features by Watching Objects Move
Explore this paper's citation graph
- Type
- article
- Published
- 2016-12-19
- Cited by
- 540
- References
- 52
- Access
- Open access
- OpenAlex
- https://openalex.org/W2575671312
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:10054272
Keywords
Computer science, Artificial intelligence, Unsupervised learning, Feature learning, Representation (politics)
References
- Unsupervised Visual Representation Learning by Context Prediction
- Learning to Segment Object Candidates
- Learning to See by Moving
- Learning to Linearize Under Uncertainty
- Fully convolutional networks for semantic segmentation
- Hypercolumns for object segmentation and fine-grained localization
- Extracting and composing robust features with denoising autoencoders
- Human action recognition by learning bases of action attributes and parts
- Learning Deep Architectures for AI
- Learning Classification with Unlabeled Data
- Robust Higher Order Potentials for Enforcing Label Consistency
- ImageNet Large Scale Visual Recognition Challenge
- SLIC Superpixels Compared to State-of-the-Art Superpixel Methods
- Principles of Object Perception
- A Unified Video Segmentation Benchmark: Annotation, Metrics and Analysis
- Structured Forests for Fast Edge Detection
- Semantic contours from inverse detectors
- Visual Parsing After Recovery From Blindness
- Histograms of oriented gradients for human detection
- ImageNet classification with deep convolutional neural networks
Cited by
- Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction
- Colorization as a Proxy Task for Visual Understanding
- Unsupervised Learning from Video to Detect Foreground Objects in Single Images
- Learning to Estimate Pose by Watching Videos
- Unsupervised Learning of Depth and Ego-Motion from Video
- Time-Contrastive Networks: Self-Supervised Learning from Multi-view Observation
- Unsupervised Learning of Object Landmarks by Factorized Spatial Embeddings
- Deep Learning Improves Template Matching by Normalized Cross Correlation
- CortexNet: a Generic Network Family for Robust Visual Temporal Representations
- Disentangling Motion, Foreground and Background Features in Videos
- Improved Speech Reconstruction from Silent Video
- Neural Expectation Maximization
- Discovery of Visual Semantics by Unsupervised and Self-Supervised Representation Learning
- Representation Learning by Learning to Count
- Multi-task Self-Supervised Visual Learning
- Lucid Data Dreaming for Multiple Object Tracking
- Playing for Benchmarks
- Pseudo-Labels for Supervised Learning on Dynamic Vision Sensor Data, Applied to Object Detection Under Ego-Motion
- Cross-Domain Self-Supervised Multi-task Feature Learning Using Synthetic Imagery
- Time-Contrastive Networks: Self-Supervised Learning from Video