The “Something Something” Video Database for Learning and Evaluating Visual Common Sense
Explore this paper's citation graph
- Type
- article
- Published
- 2017-06-13
- Cited by
- 2,046
- References
- 45
- Access
- Open access
- OpenAlex
- https://openalex.org/W2625366777
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:834612
Keywords
Computer science, Obstacle, Artificial intelligence, Commonsense reasoning, Object (grammar)
References
- Multiple View Geometry in Computer Vision
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Video (language) modeling: a baseline for generative models of natural videos
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Anticipating the future by watching unlabeled video
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
- ActivityNet: A large-scale video benchmark for human activity understanding
- Learning to Relate Images
- Large-Scale Video Classification with Convolutional Neural Networks
- A database for fine grained activity detection of cooking activities
- Unbiased look at dataset bias
- CNN Features Off-the-Shelf: An Astounding Baseline for Recognition
- A dataset for Movie Description
- Going deeper with convolutions
- Action MACH a spatio-temporal Maximum Average Correlation Height filter for action recognition
- ImageNet: A large-scale hierarchical image database
- Grounding Action Descriptions in Videos
- Vision meets robotics: The KITTI dataset
- Modeling Deep Temporal Dependencies with Recurrent "Grammar Cells"
Cited by
- ExtremeWeather: A large-scale climate dataset for semi-supervised detection, localization, and understanding of extreme weather events
- AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions
- ActivityNet Challenge 2017 Summary
- From Lifestyle Vlogs to Everyday Interactions
- Moments in Time Dataset: One Million Videos for Event Understanding
- Deep Episodic Memory: Encoding, Recalling, and Predicting Episodic Experiences for Robot Action Execution
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- Fine-grained Video Classification and Captioning
- DenseImage Network: Video Spatial-Temporal Evolution Encoding and Understanding
- Object Level Visual Reasoning in Videos
- Human Action Recognition and Prediction: A Survey
- Video benchmarks of human action datasets: a review
- Trajectory Convolution for Action Recognition
- Temporal Reasoning in Videos Using Convolutional Gated Recurrent Units
- High Order Neural Networks for Video Classification
- TSM: Temporal Shift Module for Efficient Video Understanding
- ON THE EFFECTIVENESS OF TASK GRANULARITY FOR TRANSFER LEARNING
- Perceiving Physical Equation by Observing Visual Scenarios
- Dynamic Graph Modules for Modeling Higher-Order Interactions in Activity Recognition
- Eidetic 3D LSTM: A Model for Video Prediction and Beyond
Related papers
- The Study of the Impact of Obstacle on the Efficiency of Evacuation under Different Competitive Conditions
- Efficient field courses around an obstacle
- The Analysis and Countermeasures on the Obstacle of Communicate with Human Relation in Colleges and Universities Ideological Political Work
- Efficient field courses around an obstacle
- Beating Common Sense into Interactive Applications
- Learning Common Sense through Visual Abstraction
- A viewpoint: common sense and belief formation
- Research on social common sense knowledge reasoning method based on pre-training
- On the Evaluation of Common-Sense Reasoning in Natural Language Understanding