Matrix multiplication on the Intel Touchstone Delta
Explore this paper's citation graph
Summary
An implementation is obtained that uses communication primitives highly suited to the Delta and exploits the single node assembly-coded matrix multiplication and has achieved parallel efficiency of 86 %, with overall peak performance in excess of 8 Gflops on 256 nodes for an 8800 × 8800 matrix.
- Type
- article
- Published
- 1994-10-01
- Cited by
- 51
- References
- 32
- OpenAlex
- https://openalex.org/W544946
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1691856
Keywords
Parallel computing, Computer science, FLOPS, Matrix multiplication, Scalability
References
- I860 Microprocessor Architecture
- A Parallel Implementation of the Invariant Subspace Decomposition Algorithm for Dense Symmetric Matrices
- Massively Parallel Linpack Benchmark on the Intel Touchstone Delta andIPSC/860 Systems (Progress Report)
- Communication performance of the Intel Touchstone DELTA mesh
- The Touchstone 30 Gigaflop DELTA Prototype
- Level 3 BLAS for distributed memory concurrent computers
- The impact of HPF data layout on the design of efficient and maintainable parallel linear algebra libraries
- Optimal Broadcasting in Mesh-Connected Architectures
- Solving Problems on Concurrent Processors
- Performance and Assembly Language Programming of the iPSC/860 System
- A cellular computer to implement the kalman filter algorithm
- ScaLAPACK: a scalable linear algebra library for distributed memory concurrent computers
- Multiplication of Matrices of Arbitrary Shape on a Data Parallel Computer
- Solving linear systems on vector and shared memory computers
- A set of level 3 basic linear algebra subprograms
- On parallelizable eigensolvers
- Matrix algorithms on a hypercube I: Matrix multiplication
- Comparison of scalable parallel matrix multiplication libraries
- Performance Analysis of k-Ary n-Cube Interconnection Networks
- The Multicomputer Toolbox approach to concurrent BLAS and LACS
Cited by
- Parallel Studies of the Invariant Subspace Decomposition Approach for Banded Symmetric Matrices
- Public International Benchmarks for Parallel Computers
- Public international benchmarks for parallel computers: PARKBENCH committee: Report-1
- The impact of HPF data layout on the design of efficient and maintainable parallel linear algebra libraries
- A programming model for block-structured scientific calculations on smp clusters
- Parallel Bandreduction and Tridiagonalization
- The PRISM project: infrastructure and algorithms for parallel eigensolvers
- Optimizing parallel multiplication operation for rectangular and transposed matrices
- ScaLAPACK: A Linear Algebra Library for Message-Passing Computers
- A new parallel matrix multiplication algorithm on distributed-memory concurrent computers
- Quick Matrix Multiplication on Clusters of Workstations
- A General Scalable Parallelizing of Strassen's Algorithm for Matrix Multiplication on Distributed Memory Computers
- Pumma: Parallel universal matrix multiplication algorithms on distributed memory concurrent computers
- Memory efficient parallel matrix multiplication operation for irregular problems
- Generalized Cannon's algorithm for parallel matrix multiplication
- Hierarchical approach to optimization of parallel matrix multiplication on large-scale platforms
- Parallelizing Strassen's method for matrix multiplication on distributed-memory MIMD architectures☆
- SUMMA: scalable universal matrix multiplication algorithm
- Large-scale correlated electronic structure calculations: the RI-MP2 method on parallel computers
- A high-performance matrix-multiplication algorithm on a distributed-memory parallel computer, using overlapped communication
Related papers
- Implementation of dense matrix multiplication on 2D mesh
- Parallel Matrix Multiplication: A Systematic Journey
- Matrix-vector multiplication and conjugate gradient algorithms on distributed memory computers
- The Implementation and Optimization of Parallel Linpack on Multi-Core Vector Accelerator
- Scalability of Parallel Algorithms for Matrix Multiplication
- Analysis of parallel algorithm for matrix multiplication on hypercubes
- Analysis and Comparison of Several Parallel Algorithms for Matrix Multiplication on Distributed-memory Multi-computer
- Parallel Matrix Multiplication Algorithms on Hypercube Multiprocessors
- Increasing the Efficiency of Sparse Matrix-Matrix Multiplication with a 2.5D Algorithm and One-Sided MPI