Qwen3-VL Technical Report

Explore this paper's citation graph

Summary

Qwen3-VL is introduced, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks, and three key upgrades are introduced, including an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video.

Type
preprint
Published
2025-11-26
Cited by
2,039
References
102
Access
Open access

Keywords

Security token, Key (lock), Timestamp, Code (set theory), Latency (audio)

References

Cited by

Related papers

No related papers recorded.