Evaluating Object Hallucination in Large Vision-Language Models
Explore this paper's citation graph
Summary
This work presents the first systematic study on object hallucination of LVLMs, and designs an improved evaluation method by proposing a polling-based query method called POPE, which can evaluate the object hallucinated objects in a more stable and flexible way.
- Type
- preprint
- Published
- 2023-05-17
- Cited by
- 1,922
- References
- 54
- Access
- Open access
- OpenAlex
- https://openalex.org/W4377121433
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:258740697
Keywords
Hallucinating, Object (grammar), Computer science, Visual Hallucination, Polling
References
- Im2Text: Describing Images Using 1 Million Captioned Photographs
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics (Extended Abstract)
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- VQA: Visual Question Answering
- Understanding Blind People's Experiences with Computer-Generated Captions of Social Media Images
- Neural Baby Talk
- Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
- Object Hallucination in Image Captioning
- nocaps: novel object captioning at scale
- OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
- GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
- Yin and Yang: Balancing and Answering Binary Visual Questions
- Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
- VinVL: Revisiting Visual Representations in Vision-Language Models
- Let there be a clock on the beach: Reducing Object Hallucination in Image Captioning
- A Survey of Vision-Language Pre-Trained Models
- Vision
- Computer Vision
- LAVIS: A Library for Language-Vision Intelligence
- Complexity-Based Prompting for Multi-Step Reasoning
Cited by
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
- CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
- Zero-shot Visual Question Answering with Language Model Feedback
- AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- Large Multimodal Models: Notes on CVPR 2023 Tutorial
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Visual Instruction Tuning with Polite Flamingo
- mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
- Foundational Models Defining a New Era in Vision: A Survey and Outlook
- Improving Generalization of Image Captioning with Unsupervised Prompt Learning
- Tiny LVLM-eHub: Early Multimodal Experiments with Bard
- Detecting and Preventing Hallucinations in Large Vision Language Models
- Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models
- Evaluation and Analysis of Hallucination in Large Vision-Language Models
- TouchStone: Evaluating Vision-Language Models by Language Models
- CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning
- Beyond Traditional Teaching: The Potential of Large Language Models and Chatbots in Graduate Engineering Education
- A Survey of Hallucination in Large Foundation Models
Related papers
- Auditory hallucinations in Parkinson’s disease
- The Frequency of Visual Hallucinations in Schizophrenic Patients in Saudi Arabia
- The Patient Experiences Hallucinations with Schizophrenia
- Compensatory shifts in visual perception are associated with hallucinations in Lewy body disorders
- The Strasbourg Visual Scale: A Novel Method to Assess Visual Hallucinations
- Ghosts in the machine: A Case description of visual and haptic hallucinations after right hemisphere stroke
- Visual hallucinations in the elderly.
- Occurrence and co-occurrence of hallucinations by modality in schizophrenia-spectrum disorders.
- Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training