TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Explore this paper's citation graph
Summary
TVM, a compiler that exposes graph-level and operator-level optimizations to provide performance portability to deep learning workloads across diverse hardware back-ends, and offers automated optimization of low-level programs to hardware characteristics.
- Type
- article
- Published
- 2018-10-08
- Cited by
- 2,142
- References
- 42
- OpenAlex
- https://openalex.org/W2804032941
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3296374
Keywords
Computer science, Compiler, Deep learning, Vendor, Field-programmable gate array
References
- Recurrent Neural Network Regularization
- Theano: new features and speed improvements
- BinaryConnect: Training Deep Neural Networks with binary weights during propagations
- Loo.py: transformation-based code generation for GPUs and CPUs
- Decoupled access/execute computer architectures
- Darkroom
- PuDianNao: A Polyvalent Machine Learning Accelerator
- Optimization by Simulated Annealing
- DaDianNao: A Machine-Learning Supercomputer
- Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
- Polyhedral parallel code generation for CUDA
- FFTW: an adaptive software architecture for the FFT
- Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
- Effective Hardware Based Data Prefetching for High-Performance Processors
- Automatically Tuned Linear Algebra Software
- Human-level control through deep reinforcement learning
- Simultaneous multithreading: a platform for next-generation processors
- Fast Algorithms for Convolutional Neural Networks
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- OptiML: An Implicitly Parallel Domain-Specific Language for Machine Learning
Cited by
- A Survey of FPGA-Based Neural Network Accelerator
- Demystifying differentiable programming: shift/reset the penultimate backpropagator
- Analysis of DAWNBench, a Time-to-Accuracy Machine Learning Performance Benchmark
- Deep Learning Approach for Brain Machine Interface
- Exploring the Programmability for Deep Learning Processors: from Architecture to Tensorization
- On the Anatomy of Predictive Models for Accelerating GPU Convolution Kernels and Beyond
- VTA: An Open Hardware-Software Stack for Deep Learning
- Deep learning: Computational aspects
- RLgraph: Flexible Computation Graphs for Deep Reinforcement Learning
- Meta-programming for cross-domain tensor optimizations
- FusionStitching: Deep Fusion and Code Generation for Tensorflow Computations on GPUs
- TSM: Temporal Shift Module for Efficient Video Understanding
- Dynamic Space-Time Scheduling for GPU Inference
- Towards Transparent Neural Network Acceleration
- Privacy Preserving Deep Neural Network Prediction using Trusted Hardware
- NNStreamer: Stream Processing Paradigm for Neural Networks, Toward Efficient Development and Execution of On-Device AI Applications
- Thriving in the No Man's Land between Compilers and Databases
- Hardware-conscious Query Processing in GPU-accelerated Analytical Engines
- EcoRNN: Efficient Computing of LSTM RNN Training on GPUs
- No DNN Left Behind: Improving Inference in the Cloud with Multi-Tenancy
Related papers
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Learning to Optimize Tensor Programs
- Glow: Graph Lowering Compiler Techniques for Neural Networks
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- In-datacenter performance analysis of a tensor processing unit
- TensorFlow: a system for large-scale machine learning
- Deep Residual Learning for Image Recognition