UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Explore this paper's citation graph
Summary
This work introduces UCF101 which is currently the largest dataset of human actions and provides baseline action recognition results on this new dataset using standard bag of words approach with overall performance of 44.5%.
- Type
- preprint
- Published
- 2012-12-03
- Cited by
- 7,350
- References
- 13
- Access
- Open access
- OpenAlex
- https://openalex.org/W24089286
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:7197134
Keywords
CLIPS, Upload, Computer science, Action recognition, Artificial intelligence
References
- Perceiving events and objects
- Recognizing 50 human action categories of web videos
- Actions as space-time shapes
- Recognizing realistic actions from videos “in the wild”
- Action MACH a spatio-temporal Maximum Average Correlation Height filter for action recognition
- HMDB: A large video database for human motion recognition
- Actions in context
- Action Recognition from Arbitrary Views using 3D Exemplars
- Recognizing human actions: a local SVM approach
- Detecting Carried Objects in Short Video Sequences
- Modeling Temporal Structure of Decomposable Motion Segments for Activity Classification
- Action Recognition from Arbitrary Views using 3 D Exemplars
Cited by
- Employing automatic content recognition for teaching methodology analysis in classroom videos
- Intelligent Data Engineering and Automated Learning – IDEAL 2020: 21st International Conference, Guimaraes, Portugal, November 4–6, 2020, Proceedings, Part II
- Feature sampling and partitioning for visual vocabulary generation on large action classification datasets
- The Johns Hopkins University multimodal dataset for human action recognition
- Learning Temporal Embeddings for Complex Video Analysis
- Automatic visual detection of human behavior: A review from 2000 to 2014
- Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video
- Cross Modal Distillation for Supervision Transfer
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Unsupervised Semantic Parsing of Video Collections
- Exploring Semantic Inter-Class Relationships (SIR) for Zero-Shot Action Recognition
- Supervised Learning Approaches for Automatic Structuring of Videos. (Méthodes d'apprentissage supervisé pour la structuration automatique de vidéos)
- A Robust and Efficient Video Representation for Action Recognition
- The Best of BothWorlds: Combining Data-Independent and Data-Driven Approaches for Action Recognition
- Efficient large-scale action recognition in videos using extreme learning machines
- A modified vector of locally aggregated descriptors approach for fast video classification
- Ordered trajectories for human action recognition with large number of classes
- Video Description Generation Incorporating Spatio-Temporal Features and a Soft-Attention Mechanism
- Categorizing Big Video Data on the Web: Challenges and Opportunities
- Efficient and effective human action recognition in video through motion boundary description with a compact set of trajectories
Related papers
- The Kinetics Human Action Video Dataset
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Learning realistic human actions from movies
- HMDB: A large video database for human motion recognition