BLAS on the Trident Processor: Implementation and Performance Evaluation
Explore this paper's citation graph
Summary
This paper describes the implementation of the Basic Linear Algebra Subprograms (BLAS) on the Trident processor and shows how to use the Trident parallel execution units, ring, and communication registers to effectively perform vector- vector, matrix-vector, and matrix-matrix operations needed for implementing BLAS.
- Type
- article
- Published
- 2003-01-01
- Cited by
- 0
- References
- 13
- OpenAlex
- https://openalex.org/W43909384
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:27293179
Keywords
Trident, Computer science, Parallel computing, Computer architecture, Embedded system
References
- Early 21st Century Processors - Guest Editors' Introduction
- Parallel Computers 2: Architecture, Programming and Algorithms
- Computer Architecture: A Quantitative Approach
- Basic Linear Algebra Subprograms for Fortran Usage
- A set of level 3 basic linear algebra subprograms
- An extended set of FORTRAN basic linear algebra subprograms
- The CRAY-1 computer system
- The future of wires
- Matrix computations
- Vector microprocessors
- Early 21 st Century Processors
Cited by
No citing papers recorded for this paper.
Related papers
- Parallel BLAS Performance Report
- Optimized BLAS and Its Effect on Performance of Parallel Programs
- Benchmarking Single- and Multi-Core BLAS Implementations and GPUs for use with R
- Design and Implementation of High Performance BLAS for Pentium Pro
- Performance Testing and Analysis of BLAS Libraries on Multi-Core CPUs
- GEMM-Based Level 3 BLAS: Installation, Tuning and Use of the Model Implementations and the Performance Evaluation Benchmark
- Optimization of BLAS based on Loongson 2F architecture
- Performance evaluation of multi-core intel xeon processors on basic linear algebra subprograms
- Implementation of Linear Circuit Simulator FALCON on Multi-Core CPU