Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?
Explore this paper's citation graph
- Type
- preprint
- Published
- 2017-11-27
- Cited by
- 2,263
- References
- 30
- Access
- Open access
- OpenAlex
- https://openalex.org/W2770565591
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:4539700
Keywords
Overfitting, Convolutional neural network, Computer science, Artificial intelligence, Pattern recognition (psychology)
References
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Rectified Linear Units Improve Restricted Boltzmann Machines
- ActivityNet: A large-scale video benchmark for human activity understanding
- Action recognition with trajectory-pooled deep-convolutional descriptors
- Large-Scale Video Classification with Convolutional Neural Networks
- Dense Trajectories and Motion Boundary Descriptors for Action Recognition
- Going deeper with convolutions
- ImageNet: A large-scale hierarchical image database
- HMDB: A large video database for human motion recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Deep Residual Learning for Image Recognition
- Long-Term Temporal Convolutions for Action Recognition
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Aggregated Residual Transformations for Deep Neural Networks
- Spatiotemporal Residual Networks for Video Action Recognition
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
- The Kinetics Human Action Video Dataset
Cited by
- Drive Video Analysis for the Detection of Traffic Near-Miss Incidents
- STAIR Actions: A Video Dataset of Everyday Home Actions
- Multimodal Co-Training for Selecting Good Examples from Webly Labeled Video
- Fine-grained Video Classification and Captioning
- Revisiting Temporal Modeling for Video-based Person ReID
- Segment-Tube: Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame Segmentation
- Few-Shot Adaptation for Multimedia Semantic Indexing
- Dual Viewpoint Passenger State Classification Using 3D CNNs
- Understanding human-human interactions: a survey
- Video Inpainting by Jointly Learning Temporal Structure and Spatial Details
- Bi-direction hierarchical LSTM with spatial-temporal attention for action recognition
- Activity Recognition on a Large Scale in Short Videos - Moments in Time Dataset
- Continuous Gesture Segmentation and Recognition Using 3DCNN and Convolutional LSTM
- A Dataset for Telling the Stories of Social Media Videos
- Temporal–Spatial Mapping for Action Recognition
- 3D Deep Learning from CT Scans Predicts Tumor Invasiveness of Subcentimeter Pulmonary Adenocarcinomas.
- CSI-Net: Unified Human Body Characterization and Action Recognition
- Morph: Flexible Acceleration for 3D CNN-Based Video Understanding
- Improving Human Action Recognition through Hierarchical Neural Network Classifiers
- A Large-scale RGB-D Database for Arbitrary-view Human Action Recognition
Related papers
- Challenges of Deep Learning Methods for COVID-19 Detection Using Public Datasets
- Deep Convolution Neural Network for RBC Images
- Leaf Features Extraction for Plant Classification using CNN
- Deep CNN models for pulmonary nodule classification: Model modification, model integration, and transfer learning
- Small aircraft detection using deep learning
- Automated Fruit Classification Using Deep Convolutional Neural Network
- An Improved Approach for Fire Detection using Deep Learning Models
- Avoiding Overfitting in Deep Neural Networks for Clinical Opinions Generation from General Blood Test Results