E2bird: Enhanced Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services
Explore this paper's citation graph
- Type
- article
- Published
- 2021-06-01
- Cited by
- 27
- References
- 48
- OpenAlex
- https://openalex.org/W3116103263
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:231644896
Keywords
Computer science, Quality of service, Artificial intelligence, Throughput, Deep learning
References
- Improving the speed of neural networks on CPUs
- High Performance Convolutional Neural Networks for Document Processing
- cuDNN: Efficient Primitives for Deep Learning
- Analyzing CUDA workloads using a detailed GPU simulator
- Deep learning in neural networks: An overview
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Fast Algorithms for Convolutional Neural Networks
- Poseidon: A System Architecture for Efficient GPU-based Deep Learning on Multiple Machines
- EIE: Efficient Inference Engine on Compressed Deep Neural Network
- Fixed Point Quantization of Deep Convolutional Networks
- Baymax: QoS Awareness and Increased Utilization for Non-Preemptive Accelerators in Warehouse Scale Computers
- Training Deep Nets with Sublinear Memory Cost
- TensorFlow: a system for large-scale machine learning
- Clipper: A Low-Latency Online Prediction Serving System
- S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters
- Prophet: Precise QoS Prediction on Non-Preemptive Accelerators to Improve Utilization in Warehouse-Scale Computers
- FLEP: Enabling Flexible and Efficient Preemption on GPUs
- Quality of service support for fine-grained sharing on GPUs
- QoS-Aware Scheduling of Heterogeneous Servers for Inference in Deep Neural Networks
- TensorFlow-Serving: Flexible, High-Performance ML Serving
Cited by
- Enable Simultaneous DNN Services Based on Deterministic Operator Overlap and Precise Latency Prediction
- A Survey of GPU Multitasking Methods Supported by Hardware Architecture
- Exploiting Intra-SM Parallelism in GPUs via Persistent and Elastic Blocks
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoS
- Deep Learning Workload Scheduling in GPU Datacenters: Taxonomy, Challenges and Vision
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge Servers
- EALI: Energy-aware layer-level scheduling for convolutional neural network inference services on GPUs
- BARM: A Batch-Aware Resource Manager for Boosting Multiple Neural Networks Inference on GPUs With Memory Oversubscription
- ISPA: Exploiting Intra-SM Parallelism in GPUs via Fine-Grained Resource Management
- Resource Allocation for Multiuser Edge Inference With Batching and Early Exiting
- CoFB: latency-constrained co-scheduling of flows and batches for deep learning inference service on the CPU–GPU system
- Improving Cluster Utilization Through Adaptive Resource Management for Deep Neural Network and CPU Jobs Colocation
- InferFair: Towards QoS-aware scheduling for performance isolation guarantee in heterogeneous model serving systems
- SMDP-Based Dynamic Batching for Efficient Inference on GPU-Based Platforms
- Characterizing and understanding deep neural network batching systems on GPUs
- Online Resource Provisioning and Batch Scheduling for AIoT Inference Serving in an XPU Edge Cloud
- Serving DNN Inference With Fine-Grained Spatio-Temporal Sharing of GPU Servers
- SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
- DynMap: A Heuristic Dynamic Mapper for CGRA Multitask Dynamic Resource Allocation
- Over-the-Air Edge Inference for Low-Altitude Airspace: Generative AI-Aided Multi-Task Batching and Beamforming Design
Related papers
- Deep fake Detection Through Deep Learning
- Evaluation of AHS throughput using SmartCap
- Study of the throughput of WLAN according to channeling and density of AP
- POSITIVE-NEGATIVE ASYMMETRY IN MENTAL STATE INFERENCE: REPLICATION AND EXTENSION
- Extending Little’s Law to single order throughput times
- Why & When Deep Learning Works: Looking Inside Deep Learnings
- A Note on Latency Variability of Deep Neural Networks for Mobile Inference
- On the Latency Variability of Deep Neural Networks for Mobile Inference
- Empowering Adaptive Early-Exit Inference with Latency Awareness
- CryptoNAS: Private Inference on a ReLU Budget