DeViSE: A Deep Visual-Semantic Embedding Model
Explore this paper's citation graph
Summary
This paper presents a new deep visual-semantic embedding model trained to identify visual objects using both labeled image data as well as semantic information gleaned from unannotated text and shows that the semantic information can be exploited to make predictions about tens of thousands of image labels not observed during training.
- Type
- article
- Published
- 2013-12-05
- Cited by
- 3,069
- References
- 22
- OpenAlex
- https://openalex.org/W2123024445
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:261138
Keywords
Computer science, Leverage (statistics), Embedding, Artificial intelligence, Class (philosophy)
References
- Hierarchical Probabilistic Neural Network Language Model
- Efficient Estimation of Word Representations in Vector Space
- Improving neural networks by preventing co-adaptation of feature detectors
- Evaluating knowledge transfer and zero-shot learning in a large-scale setting
- Large scale image annotation: learning to rank with joint word-image embeddings
- ImageNet: A large-scale hierarchical image database
- ImageNet Large Scale Visual Recognition Challenge
- Fast, Accurate Detection of 100,000 Object Classes on a Single Machine
- Zero-Shot Learning Through Cross-Modal Transfer
- Zero-shot Learning with Semantic Output Codes
- Distributed Representations of Words and Phrases and their Compositionality
- Label Embedding Trees for Large Multi-Class Tasks
- ImageNet classification with deep convolutional neural networks
- Large Scale Distributed Deep Networks
- Visual and semantic similarity in ImageNet
- The Importance of Encoding Versus Training with Sparse Coding and Vector Quantization
- Visualizing Data using t-SNE
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
- A Neural Probabilistic Language Model
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Cited by
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Learning and Recognizing The Hierarchical and Sequential Structure of Human Activities
- Weakly-Supervised Alignment of Video with Text
- Learning Temporal Embeddings for Complex Video Analysis
- SHOE: Supervised Hashing with Output Embeddings
- Semantically enabled image similarity search
- Learning Contextualized Semantics from Co-occurring Terms via a Siamese Architecture
- Cross Modal Distillation for Supervision Transfer
- Learning Contextualized Music Semantics from Tags via a Siamese Network
- Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified Framework
- Modality-Dependent Cross-Media Retrieval
- Recognizing unknown objects with attributes relationship model
- Effective deep learning-based multi-modal retrieval
- Combining Language and Vision with a Multimodal Skip-gram Model
- Semantic Graph for Zero-Shot Learning
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Modeling Documents with Event Model
- Joint Learning of Distributed Representations for Images and Texts
- Deep learning helicopter dynamics models
- Mining Latent Attributes From Click-Through Logs for Image Recognition
Related papers
- Zero-Shot Learning by Convex Combination of Semantic Embeddings
- Evaluation of Output Embeddings for Fine-grained Image Classification
- Learning a Deep Embedding Model for Zero-Shot Learning
- Latent Embeddings for Zero-Shot Classification
- Synthesized Classifiers for Zero-Shot Learning
- GloVe: Global Vectors for Word Representation
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks