Matrix multiplication on the Intel Touchstone Delta

Explore this paper's citation graph

Summary

An implementation is obtained that uses communication primitives highly suited to the Delta and exploits the single node assembly-coded matrix multiplication and has achieved parallel efficiency of 86 %, with overall peak performance in excess of 8 Gflops on 256 nodes for an 8800 × 8800 matrix.

Type
article
Published
1994-10-01
Cited by
51
References
32

Keywords

Parallel computing, Computer science, FLOPS, Matrix multiplication, Scalability

References

Cited by

Related papers