Spatio-temporal deformable 3D ConvNets with attention for action recognition
Explore this paper's citation graph
Summary
This paper proposes a spatio-temporal deformable ConvNet module with an attention mechanism, which takes into consideration the mutual correlations in both temporal and spatial domains, to effectively capture the long-range and long-distance dependencies in the video actions.
- Type
- article
- Published
- 2020-02-01
- Cited by
- 131
- References
- 61
- OpenAlex
- https://openalex.org/W2971915722
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:203033155
Keywords
Computer science, Action recognition, Artificial intelligence, Convolution (computer science), Motion (physics)
References
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Fully convolutional networks for semantic segmentation
- Beyond short snippets: Deep networks for video classification
- Action recognition with trajectory-pooled deep-convolutional descriptors
- On Space-Time Interest Points
- A non-local algorithm for image denoising
- A 3-dimensional sift descriptor and its application to action recognition
- Action recognition by dense trajectories
- HMDB: A large video database for human motion recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Compressive Sequential Learning for Action Similarity Labeling
- Do less and achieve more: Training CNNs for action recognition utilizing action images from the Web
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Saliency-Aware Video Object Segmentation
- Deformable Convolutional Networks
- Active Convolution: Learning the Shape of Convolution for Image Classification
- Two Stream LSTM: A Deep Fusion Framework for Human Action Recognition
Cited by
- Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action Classification
- Gated CNN: Integrating multi-scale feature layers for object detection
- Spectral rotation for deep one-step clustering
- Heterogenous output regression network for direct face alignment
- Projection based weight normalization: Efficient method for optimization on oblique manifold in DNNs
- Adaptive iterative attack towards explainable adversarial robustness
- Binary Neural Networks: A Survey
- Fast Nearest Subspace Search via Random Angular Hashing
- Deep quantization generative networks
- Action Recognition in Videos Using Pre-Trained 2D Convolutional Neural Networks
- Timed-image based deep learning for action recognition in video sequences
- Deep transductive network for generalized zero shot learning
- Unsupervised urban scene segmentation via domain adaptation
- Recurrent bag-of-features for visual information analysis
- A Context Based Deep Temporal Embedding Network in Action Recognition
- Self-attention driven adversarial similarity learning network
- Candidate region correlation for video action detection
- Lightweight dynamic conditional GAN with pyramid attention for text-to-image synthesis
- Towards Non-I.I.D. image classification: A dataset and baselines
- Play and rewind: Context-aware video temporal action proposals
Related papers
- Balanced group convolution: an improved group convolution based on approximability estimates
- The Urbanik generalized convolutions in the non-commutative probability and a forgotten method of constructing generalized convolution
- The conjectures of Embrechts and Goldie
- Human-like action recognition system on whole body motion-captured file
- MultiAct: Long-Term 3D Human Motion Generation from Multiple Action Labels
- STSM: Spatio-Temporal Shift Module for Efficient Action Recognition
- MultiAct: Long-Term 3D Human Motion Generation from Multiple Action Labels
- Research on Human Action Recognition Based on Convolutional Neural Network
- Human-Like Daily Action Recognition Model
- Driving Action Recognition Based on 3D Convolution