nocaps: novel object captioning at scale
Explore this paper's citation graph
- Type
- article
- Published
- 2019-10-01
- Cited by
- 671
- References
- 54
- Access
- Open access
- OpenAlex
- https://openalex.org/W2904565150
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:56517630
Keywords
Closed captioning, Benchmark (surveying), Object (grammar), Object detection, Image (mathematics)
References
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- CIDEr: Consensus-based image description evaluation
- BabyTalk: Understanding and Generating Simple Image Descriptions
- Maximum likelihood from incomplete data via the EM - algorithm plus discussions on the paper
- Bleu: a Method for Automatic Evaluation of Machine Translation
- ImageNet: A large-scale hierarchical image database
- Im2Text: Describing Images Using 1 Million Captioned Photographs
- ImageNet Large Scale Visual Recognition Challenge
- METEOR: An Automatic Metric for MT Evaluation with High Levels of Correlation with Human Judgments
- ROUGE: A Package for Automatic Evaluation of Summaries
- Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data
- From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
- Visualizing Data using t-SNE
Cited by
- A Systematic Literature Review on Image Captioning
- Know What You Don’t Know: Modeling a Pragmatic Speaker that Refers to Objects of Unknown Categories
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
- Visual Understanding through Natural Language
- Compositional Generalization in Image Captioning
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Meshed-Memory Transformer for Image Captioning
- Understanding Image Captioning Models beyond Visualizing Attention
- Captioning Images Taken by People Who Are Blind
- Captioning Images with Novel Objects via Online Vocabulary Expansion
- TextCaps: a Dataset for Image Captioning with Reading Comprehension
- Grounded Situation Recognition
- Egoshots, an ego-vision life-logging dataset and semantic fidelity metric to evaluate diversity in image captioning models
- Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
- CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
- VirTex: Learning Visual Representations from Textual Annotations
- Say As You Wish: Fine-Grained Control of Image Caption Generation With Abstract Scene Graphs
- Image Captioning menurut Scientific Revolution Kuhn dan Popper
- VIVO: Surpassing Human Performance in Novel Object Captioning with Visual Vocabulary Pre-Training
Related papers
No related papers recorded.