Qwen3-VL Technical Report
Explore this paper's citation graph
Summary
Qwen3-VL is introduced, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks, and three key upgrades are introduced, including an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video.
- Type
- preprint
- Published
- 2025-11-26
- Cited by
- 2,039
- References
- 102
- Access
- Open access
- OpenAlex
- https://openalex.org/W4416780398
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:283262018
Keywords
Security token, Key (lock), Timestamp, Code (set theory), Latency (audio)
References
- SUN RGB-D: A RGB-D scene understanding benchmark suite
- Generation and Comprehension of Unambiguous Object Descriptions
- ReferItGame: Referring to Objects in Photographs of Natural Scenes
- Billion-Scale Similarity Search with GPUs
- TALL: Temporal Activity Localization via Language Query
- A Diagram is Worth a Dozen Images
- Objects365: A Large-Scale, High-Quality Dataset for Object Detection
- DocVQA: A Dataset for VQA on Document Images
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- InfographicVQA
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Grounded Language-Image Pre-training
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved With Text
- OBELISC: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
- Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Instruction-Following Evaluation for Large Language Models
- EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
- Teaching CLIP to Count to Ten
Cited by
- A Survey on Deep Learning Techniques for Action Anticipation
- PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
- Modality-Inconsistent Continual Learning of Multimodal Large Language Models
- Multi-P2A: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models
- BioMiner: A Multi-modal System for Automated Mining of Protein-Ligand Bioactivity Data from Literature
- FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
- RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
- Mixture of Experts in Large Language Models
- HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
- SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets
- ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
Related papers
No related papers recorded.