Multimodal Deep Learning
Explore this paper's citation graph
Summary
This work presents a series of tasks for multimodal learning and shows how to train deep networks that learn features to address these tasks, and demonstrates cross modality feature learning, where better features for one modality can be learned if multiple modalities are present at feature learning time.
- Type
- article
- Published
- 2011-06-28
- Cited by
- 3,593
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W2184188583
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:352650
Keywords
Computer science, Modalities, Feature learning, Artificial intelligence, Deep learning
References
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- See me, hear me: integrating automatic speech recognition and lip-reading
- Adaptive multimodal fusion by uncertainty compensation
- The challenge of multispeaker lip-reading
- Hearing lips and seeing voices
- Lipreading and audio-visual speech perception.
- Extracting and composing robust features with denoising autoencoders
- CUAVE: A new audio-visual database for multimodal human-computer interface research
- Canonical Correlation Analysis: An Overview with Application to Learning Methods
- Extraction of Visual Features for Lipreading
- Training Products of Experts by Minimizing Contrastive Divergence
- Self-taught learning: transfer learning from unlabeled data
- Information Theoretic Feature Extraction for Audio-Visual Speech Recognition
- Sparse deep belief net model for visual area V2
- A Fast Learning Algorithm for Deep Belief Nets
- Adaptive Multimodal Fusion by Uncertainty Compensation With Application to Audiovisual Speech Recognition
- Integration of acoustic and visual speech signals using neural networks
- Adaptive bimodal sensor fusion for automatic speechreading
- "Eigenlips" for robust speech recognition
- Histograms of oriented gradients for human detection
Cited by
- A multiple kernel classification approach based on a Quadratic Successive Geometric Segmentation methodology with a fault diagnosis case.
- Stacked Denoising Tensor Auto-Encoder for Action Recognition With Spatiotemporal Corruptions
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Not all signals are created equal : Dynamic objective auto-encoder for multivariate data
- An Overview of Deep-Structured Learning for Information Processing
- Multimodal learning with deep Boltzmann machines
- Learning joint representation for community question answering with tri-modal DBM
- Online Multi-Stage Deep Architectures for Feature Extraction and Object Recognition
- Apprentissage de représentations et robotique développementale : quelques apports de l'apprentissage profond pour la robotique autonome. (Representation learning and developmental robotics : on the use of deep learning for autonomous robots)
- Modeling the Variability in Brain Morphology and Lesion Distribution in Multiple Sclerosis by Deep Learning
- Learning Representations with a Dynamic Objective Sparse Autoencoder
- Combining heterogeneous deep neural networks with conditional random fields for Chinese dialogue act recognition
- Learning Multiple Tasks with Multilinear Relationship Networks
- Learning Deep Representations : Toward a better new understanding of the deep learning paradigm. (Apprentissage de représentations profondes : vers une meilleure compréhension du paradigme d'apprentissage profond)
- Cross Modal Distillation for Supervision Transfer
- Correlational Neural Networks
- Integrating articulatory data in deep neural network-based acoustic modeling
- Big Data Analytics in Bioinformatics: A Machine Learning Perspective
- Effective deep learning-based multi-modal retrieval
- The Dynamic Brain: Modeling Neural Dynamics and Interactions From Human Electrophysiological Recordings
Related papers
- Deep Learning
- Multimodal Machine Learning: A Survey and Taxonomy
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks
- A Fast Learning Algorithm for Deep Belief Nets
- Gradient-based learning applied to document recognition
- ImageNet: A large-scale hierarchical image database
- A new approach to cross-modal multimedia retrieval