Srift: Swift and Thrift Cloud-Based Distributed Training
Explore this paper's citation graph
Summary
It is shown Srift's choices of VM instances can lead to up to 2x better throughput and 1.6x lower cost per iteration compared to baseline choices across various DNN models in real-world scenarios, leveraging heterogeneous setups and spot instances.
- Type
- preprint
- Published
- 2020-11-29
- Cited by
- 1
- References
- 79
- Access
- Open access
- OpenAlex
- https://openalex.org/W3108415931
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:227227725
Keywords
Throughput, Computer science, Cloud computing, Baseline (sea), Training (meteorology)
References
- DimmWitted: A Study of Main-Memory Statistical Analytics
- Architectures and message-passing algorithms for cluster computing: Design and performance
- Inside the Social Network's (Datacenter) Network
- A Comparative Study Of Data Center Network Architectures
- Prediction of TCP throughput: formula-based and history-based methods
- An architecture for parallel topic models
- High performance network virtualization with SR-IOV
- Scaling Distributed Machine Learning with the Parameter Server
- PortLand: a scalable fault-tolerant layer 2 data center network fabric
- Communication Efficient Distributed Machine Learning with the Parameter Server
- Practical Bayesian Optimization of Machine Learning Algorithms
- Optimization of Collective Communication Operations in MPICH
- VL2: a scalable and flexible data center network
- Deep Residual Learning for Image Recognition
- Scalable collective message-passing algorithms
- On robust estimation of the location parameter
- GeePS: scalable deep learning on distributed GPUs with a GPU-specialized parameter server
- Clipper: A Low-Latency Online Prediction Serving System
- IncBricks: Toward In-Network Computation with an In-Network Cache
- In-datacenter performance analysis of a tensor processing unit
Cited by
Related papers
- Ernest: Efficient Performance Prediction for Large-Scale Advanced Analytics
- Task Balanced Workflow Scheduling Technique considering Task Processing Rate in Spot Market
- Uncertainty-Aware Elastic Virtual Machine Scheduling for Stream Processing Systems
- Scheduling for Energy Efficiency and Throughput Maximization in a Faulty Cloud Environment
- LIBRA: Enabling Workload-Aware Multi-Dimensional Network Topology Optimization for Distributed Training of Large AI Models
- Lynceus: Cost-efficient Tuning and Provisioning of Data Analytic Jobs
- Cost-effective Resource Provisioning for Spark Workloads
- Dynamic provisioning with structure inspired selection and limitation of VMs based cost-time efficient workflow scheduling in the cloud