MiMatrix: A Massively Distributed Deep Learning Framework on a Petascale High-density Heterogeneous Cluster
Explore this paper's citation graph
Summary
This paper develops and implements a novel job server parallel software framework, named by "MiMatrix", for distributed deep learning training, and proposes a novel GPUDirect Remote direct memory access~(RDMA)-aware parallel algorithm of AllReucde executed by computing servers.
- Type
- article
- Published
- 2018-02-07
- Cited by
- 0
- References
- 42
- Access
- Open access
- OpenAlex
- https://openalex.org/W2787290722
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:195346563
Keywords
Computer science, Petascale computing, Bottleneck, Distributed computing, Server
References
- Project Adam: Building an Efficient and Scalable Deep Learning Training System
- Computer Architecture: A Quantitative Approach
- MPI: A Message-Passing Interface Standard
- cuDNN: Efficient Primitives for Deep Learning
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Big Data Deep Learning: Challenges and Perspectives
- Learning Deep Architectures for AI
- Deep learning in neural networks: An overview
- A unified architecture for natural language processing: deep neural networks with multitask learning
- ImageNet Large Scale Visual Recognition Challenge
- Petuum: A New Platform for Distributed Machine Learning on Big Data
- Deep learning with COTS HPC systems
- Large Scale Distributed Deep Networks
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
- Deep Residual Learning for Image Recognition
- Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning
- Revisiting Distributed Synchronous SGD
- TensorFlow: a system for large-scale machine learning
Cited by
No citing papers recorded for this paper.
Related papers
- A Novel Co-design Peta-scale Heterogeneous Cluster for Deep Learning Training
- Optimizing execution for pipelined‐based distributed deep learning in a heterogeneously networked GPU cluster
- ShmCaffe: A Distributed Deep Learning Platform with Shared Memory Buffer for HPC Architecture
- A multi-agent architecture for scheduling of high performance services in a GPU cluster
- Advanced environments, tools, and applications for cluster computing : NATO advanced research workshop, IWCC 2001 : Mangalia, Romania, September 1-6, 2001 : revised papers
- KernelHive: a new workflow‐based framework for multilevel high performance computing using clusters and workstations with CPUs and GPUs
- A virtual memory based runtime to support multi-tenancy in clusters with GPUs
- Optimising MPI Applications for Heterogeneous Coupled Clusters with MetaMPICH