Vision Transformers for Dense Prediction
Explore this paper's citation graph
- Type
- preprint
- Published
- 2021-03-24
- Cited by
- 3,022
- References
- 60
- Access
- Open access
- OpenAlex
- https://openalex.org/W3136635488
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:232352612
Keywords
Computer science, Transformer, Artificial intelligence, Architecture, Pascal (unit)
References
- Learning Deconvolution Network for Semantic Segmentation
- Fully convolutional networks for semantic segmentation
- SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation
- ImageNet: A large-scale hierarchical image database
- The Role of Context for Object Detection and Semantic Segmentation in the Wild
- Efficient Visual Search of Videos Cast as Text Retrieval
- Are we ready for autonomous driving? The KITTI vision benchmark suite
- Deep Residual Learning for Image Recognition
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- Pyramid Scene Parsing Network
- RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Scene Parsing through ADE20K Dataset
- A Multi-view Stereo Benchmark with High-Resolution Images and Multi-camera Videos
- Monocular Relative Depth Perception with Web Stereo Data Supervision
- Deep Ordinal Regression Network for Monocular Depth Estimation
- Exploring the Limits of Weakly Supervised Pretraining
- CornerNet: Detecting Objects as Paired Keypoints
Cited by
- Deep-learning-based pyramid-transformer for localized porosity analysis of hot-press sintered ceramic paste
- SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark Under Multiple Environments
- Monocular Depth Estimation Through Virtual-World Supervision and Real-World SfM Self-Supervision
- Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian Loss
- From Noon to Sunset: Interactive Rendering, Relighting, and Recolouring of Landscape Photographs by Modifying Solar Position
- Lightweight Monocular Depth with a Novel Neural Architecture Search Method
- Light Field Image Super-Resolution With Transformers
- Scaled ReLU Matters for Training Vision Transformers
- Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-supervised Learning
- D-Net: A Generalised and Optimised Deep Network for Monocular Depth Estimation
- Vision Transformer Hashing for Image Retrieval
- Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans
- Multi-Task Self-Training for Learning General Representations
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Learning multiplane images from single views with self-supervision
- Body Size and Depth Disambiguation in Multi-Person Reconstruction from Single Images
- A Review of Benchmark Datasets and Training Loss Functions in Neural Depth Estimation
- ConvNets vs. Transformers: Whose Visual Representations are More Transferable?
- Monocular Human Depth Estimation Via Pose Estimation
- SwinNet: Swin Transformer Drives Edge-Aware RGB-D and RGB-T Salient Object Detection