Optimization Techniques for Mapping Algorithms and Applications onto CUDA GPU Platforms and CPU-GPU Heterogeneous Platforms
Explore this paper's citation graph
Summary
A highly multithreaded FFT-based direct Poisson solver that is optimized for the recent NVIDIA GPUs and a new approach that minimizes the number of global memory accesses and overlaps the computations along the different dimensions are presented.
- Published
- 2014-01-01
- Cited by
- 0
- References
- 81
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:22009949
References
- Using OpenMP: Portable Shared Memory Parallel Programming (Scientific and Engineering Computation)
- A Fundamental Turn Toward Concurrency in Software
- Algorithmic and software challenges when moving towards exascale
- Computer Architecture, Fifth Edition: A Quantitative Approach
- Parallel Programming
- A direct Method for the Discrete Solution of Separable Elliptic Equations
- AMD 3DNow! technology: architecture and implementations
- Large-scale FFT on GPU clusters
- Cyclic Reduction Tridiagonal Solvers on GPUs Applied to Mixed-Precision Multigrid
- The uniform memory hierarchy model of computation
- OpenCL: A Parallel Programming Standard for Heterogeneous Computing Systems
- Scalable GPU graph traversal
- PyCUDA and PyOpenCL: A scripting-based approach to GPU run-time code generation
- Fast and Accurate Simulation of the Cray XMT Multithreaded Supercomputer
- ArrayFire: a GPU acceleration platform
- A Fast Direct Solution of Poisson's Equation Using Fourier Analysis
- An empirically tuned 2D and 3D FFT library on CUDA GPU
- Optimized strategies for mapping three-dimensional FFTs onto CUDA GPUs
- THE METHODS OF CYCLIC REDUCTION, FOURIER ANALYSIS AND THE FACR ALGORITHM FOR THE DISCRETE SOLUTION OF POISSON'S EQUATION ON A RECTANGLE*
- clSpMV: A Cross-Platform OpenCL SpMV Framework on GPUs
Cited by
No citing papers recorded for this paper.
Related papers
No related papers recorded.