YOLOv12: Attention-Centric Real-Time Object Detectors
Explore this paper's citation graph
Summary
This paper proposes an attention-centric YOLO framework, namely YOLOv12, that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms.
- Type
- preprint
- Published
- 2025-02-18
- Cited by
- 2,318
- References
- 79
- Access
- Open access
- OpenAlex
- https://openalex.org/W4407759488
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:276422323
Keywords
Computer science, Object (grammar), Detector, Computer vision, Artificial intelligence
References
- You Only Look Once: Unified, Real-Time Object Detection
- YOLO9000: Better, Faster, Stronger
- YOLOv3: An Incremental Improvement
- Albumentations: fast and flexible image augmentations
- Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- mixup: Beyond Empirical Risk Minimization
- CCNet: Criss-Cross Attention for Semantic Segmentation
- IoU Loss for 2D/3D Object Detection
- Mobile Robot Navigation Using an Object Recognition Software with RGBD Images and the YOLO Algorithm
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- CSPNet: A New Backbone that can Enhance Learning Capability of CNN
- Efficient Attention: Attention with Linear Complexities
- Axial Attention in Multidimensional Transformers
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- AP-Loss for Accurate One-Stage Object Detection
- End-to-End Object Detection with Transformers
- Linformer: Self-Attention with Linear Complexity
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- AutoAssign: Differentiable Label Assignment for Dense Object Detection
Cited by
- RBF Weighted Hyper-Involution for RGB-D Object Detection
- 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Views
- Comprehensive Performance Evaluation of YOLOv12, YOLO11, YOLOv10, YOLOv9 and YOLOv8 on Detecting and Counting Fruitlet in Complex Orchard Environments
- Aerial Flood Scene Classification Using Fine-Tuned Attention-Based Architecture for Flood-Prone Countries in South Asia
- RSNet: A Light Framework for The Detection of SAR Ship Detection
- Multiple Object Detection and Tracking in Panoramic Videos for Cycling Safety Analysis
- Error Slice Discovery via Manifold Compactness
- MHAF-YOLO: Multi-Branch Heterogeneous Auxiliary Fusion YOLO for accurate object detection
- Citrus Disease Detection Based on Dilated Reparam Feature Enhancement and Shared Parameter Head
- A real-time insulator condition detection model for UAV inspection based on FG-YOLO
- Comparative Analysis of Deep Neural Networks YOLOv11 and YOLOv12 for Real-Time Vehicle Detection in Autonomous Vehicles
- Deep Learning-Based NIR Face Detection under Adverse Illumination with Explainable AI
- MAF-MixNet: Few-Shot Tea Disease Detection Based on Mixed Attention and Multi-Path Feature Fusion
- Performance Analysis of Glass Surface Detection Based on the YOLOs
- OE-YOLO: An EfficientNet-Based YOLO Network for Rice Panicle Detection
- MC-ASFF-ShipYOLO: Improved Algorithm for Small-Target and Multi-Scale Ship Detection for Synthetic Aperture Radar (SAR) Images
- MERS-Net: A Lightweight and Efficient Remote Sensing Image Object Detector
- CPDD: A Cross-Scenario Photovoltaic Defect Detector Based on Fine-Grained Feature Autoencoding and Pseudo-Box Contrastive Learning
- CSSA-YOLO: Cross-Scale Spatiotemporal Attention Network for Fine-Grained Behavior Recognition in Classroom Environments
- Visual large language models for welding assessment
Related papers
- An Object Detection and Pose Estimation Approach for Position Based Visual Servoing
- Self-monitoring to improve robustness of 3D object tracking for robotics
- Foreground object segmentation from binocular stereo video
- 6-DOF object localization by combining monocular vision and robot arm kinematics
- Tracking in 3D: Image Variability Decomposition for Recovering Object Pose and Illumination
- Hand-eye calibration using a single image and robotic picking up using images lacking in contrast
- Motion-based Object Detection and Tracking in Color Image Sequence
- Transparent object detection and location based on RGB-D camera
- Stereo object tracking with fusion of texture, color and disparity information
- Robust object tracking based on RGB-D camera