Software Libraries for Linear Algebra Computations on High Performance Computers
Explore this paper's citation graph
Summary
This paper discusses the design of linear algebra libraries for high performance computers, with particular emphasis on the development of scalable algorithms for multiple instruction multiple data (MIMD) distributed memory concurrent computers.
- Type
- article
- Published
- 1995-06-01
- Cited by
- 142
- References
- 56
- OpenAlex
- https://openalex.org/W2012309011
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3167923
Keywords
Computer science, Cholesky decomposition, Parallel computing, Linear algebra, Distributed memory
References
- LAPACK Working Note 16: `Results from the Initial Release of LAPACK,''
- LAPPACK Working Note No. 28: The IBM RISC System/6000 and Linear Algebra Operations
- Broadcasting on linear arrays and meshes
- An object oriented design for high performance linear algebra on distributed memory architectures
- Communication performance of the Intel Touchstone DELTA mesh
- LAPACK for Distributed Memory Architectures: Progress Report
- Matrix factorization on a hypercube multiprocessor
- Increasing the performance of mathematical software through high-level modularity
- Two Dimensional Basic Linear Algebra Communication Subprograms
- Solving Problems on Concurrent Processors
- On the scalability of FFT on parallel computers
- A Proposal for a User-Level, Message-Passing Interface in a Distributed Memory Environment
- The design of scalable software libraries for distributed memory concurrent computers
- ScaLAPACK: a scalable linear algebra library for distributed memory concurrent computers
- Data redistribution and concurrency
- Basic Linear Algebra Subprograms for Fortran Usage
- Solving linear systems on vector and shared memory computers
- A set of level 3 basic linear algebra subprograms
- Handbook for Automatic Computation. Vol II, Linear Algebra
- Large Dense Numerical Linear Algebra in 1993: the Parallel Computing Influence
Cited by
- Nonlinear kernel-based statistical pattern analysis
- A parameterized ordering for cache-, register- and pipeline-efficient Givens QR decomposition
- Loop partitioning versus tiling for cache-based multiprocessors
- Contributions à la recherche en calcul scientifique haute performance pour les matrices creuses
- A Parallel Version of the Unsymmetric Lanczos Algorithm and its Application to QMR
- Tiling for Heterogeneous Computing Platforms
- A Proposal for a Heterogeneous Cluster ScaLAPACK (Dense Linear Solvers)
- Parametric Micro-level Performance Models for Parallel Computing
- Performance Analysis of Cache-Aware Multicore Parallelization with Application to Optimization Theory
- Algorithmic and Scheduling Techniques for Heterogeneous and Distributed Computing
- Efficient LU Factorization for Texas Instruments Keystone Architecture Digital Signal Processors
- Determining the idle time of a tiling: new results
- Data Allocation Strategies for Dense Linear Algebra Kernels on Heterogeneous Two-dimensional Grids
- Final Report, Center for Programming Models for Scalable Parallel Computing: Co-Array Fortran, Grant Number DE-FC02-01ER25505
- Filmification of Methods: Computation on Matrices
- Real time resource management and adaptive parallel programming for a cluster of computers: a comparison of different approaches in a computationally intensive environment
- ScaLAPACK: A Linear Algebra Library for Message-Passing Computers
- Parallel Computer Architecture: A Hardware/Software Approach
- Load balancing strategies for dense linear algebra kernels on heterogeneous two-dimensional grids
- Programming language requirements for the next millennium