Distributed TensorFlow with MPI
Explore this paper's citation graph
Summary
This paper extends recently proposed Google TensorFlow for execution on large scale clusters using Message Passing Interface (MPI) and requires minimal changes to the Tensorflow runtime -- making the proposed implementation generic and readily usable to increasingly large users of Tensor Flow.
- Type
- preprint
- Published
- 2016-03-07
- Cited by
- 40
- References
- 16
- Access
- Open access
- OpenAlex
- https://openalex.org/W2294581108
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:2295647
Keywords
Computer science, Distributed computing
References
- A High-Performance, Portable Implementation of the MPI Message Passing Interface Standard
- Going deeper with convolutions
- Searching for exotic particles in high-energy physics with deep learning
- The WEKA data mining software: an update
- Searching for Higgs Boson Decay Modes with Deep Learning
- Machine Learning and Its Applications to Biology
- LIBSVM: A library for support vector machines
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Scientific Discovery at the Exascale
- A Comparative Study of Classification Techniques On Adult Data Set
- Scikit-learn: Machine Learning in Python
- Synergistic Challenges in Data-Intensive Science and Exascale Computing
- Support Vector Machines in High Energy Physics
- MPI-2: Extending the Message-Passing Interface
Cited by
- S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters
- User-transparent Distributed TensorFlow
- Building TensorFlow Applications in Smart City Scenarios
- Accelerating High-energy Physics Exploration with Deep Learning
- What does fault tolerant deep learning need from MPI?
- Improving the Performance of Distributed TensorFlow with RDMA
- Characterizing Deep Learning over Big Data (DLoBD) Stacks on RDMA-Capable Networks
- Distributed Learning of CNNs on Heterogeneous CPU/GPU Architectures
- RLlib: Abstractions for Distributed Reinforcement Learning
- Towards Scalable Deep Learning via I/O Analysis and Optimization
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Parallel I/O Optimizations for Scalable Deep Learning
- Gear Training: A new way to implement high-performance model-parallel training
- NUMA-Caffe
- DLoBD: A Comprehensive Study of Deep Learning over Big Data Stacks on HPC Clusters
- SingleCaffe: An Efficient Framework for Deep Learning on a Single Node
- STEP : A Distributed Multi-threading Framework Towards Efficient Data Analytics
- Accelerating TensorFlow with Adaptive RDMA-Based gRPC
- Various Frameworks and Libraries of Machine Learning and Deep Learning: A Survey
- Scalable Deep Learning on Distributed Infrastructures
Related papers
- ИСПОЛЬЗОВAНИЕ ПОТЕНЦИAЛA СОЦИAЛЬНЫХ ПAРТНЕРОВ В ПОДГОТОВКЕ БУДУЩИХ ПЕДAГОГОВ
- Using DataGrid Control to Realize DataBase of Querying in VB6.0
- Study and Two Types of Typical Usage of DataGrid Web Server Control
- PACWON: A parallelizing compiler for workstations on a network
- Bidirectional Sort and Choosing a Row to Update or Delete by Click Any Cell in DataGrid
- OpenCL-accelerated object classification in video streams using Spatial Pooler of Hierarchical Temporal Memory
- ESKVS: efficient and secure approach for keyframes-based video summarization framework
- Flexible Application of VSFlexGrid