Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures
Explore this paper's citation graph
- Type
- report
- Published
- 2008-01-01
- Cited by
- 1,814
- References
- 37
- Access
- Open access
- OpenAlex
- https://openalex.org/W4214826206
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5703612
Keywords
Transpose, Computer science, Multi-core processor, Parallel computing, Code (set theory)
References
- Analyzing the behavior and performance of parallel programs
- Auto-tuning performance on multicore computers
- Lattice Boltzmann simulation optimization on leading multicore platforms
- Latency Lags Bandwidth
- Computer Architecture: A Quantitative Approach
- Performance of Synchronized Iterative Processes in Multiprocessor Systems
- Analytic Queueing Network Models for Parallel Processing of Task Systems
- Amdahl’s Law in the multicore era
- Evaluating Associativity in CPU Caches
- Mapping computational concepts to GPUs
- A Hierarchical Approach to Modeling and Improving the Performance of Scientific Applications on the KSR1
- The Design and Implementation of FFTW3
- Performance Optimizations and Bounds for Sparse Matrix-Vector Multiply
- A genetic algorithms approach to modeling the performance of memory-bound computations
- The Landscape of Parallel Computing Research: A View from Berkeley
- Improving the ratio of memory operations to floating-point operations in loops
- Self-Adapting Linear Algebra Algorithms and Software
- Validity of the single processor approach to achieving large scale computing capabilities
- Stencil computation optimization and auto-tuning on state-of-the-art multicore architectures
- The PARSEC benchmark suite: Characterization and architectural implications
Cited by
- Orientation and rotational diffusion of fibers in semidilute suspension
- Power Management for GPU-CPU Heterogeneous Systems
- Towards Real-time SAR
- Exploring power efficiency and optimizations targeting heterogeneous applications
- The parallelization of binary decision diagram operations for model checking
- Optimizing for a Many-Core Architecture without Compromising Ease-of-Programming
- Performance analysis and fitness of GPGPU and multicore architectures for scientific applications
- Chip‐level and multi‐node analysis of energy‐optimized lattice Boltzmann CFD simulations
- Connecting architecture, fitness, optimizations and performance using an anisotropic diffusion filter
- An OpenCL-based Solution for Portable Bodyscan SAR Processing on Multicore Platforms
- Scalable multi-core model checking
- Performance Analysis of Hybrid CPU/GPU Environments
- DiamondTorre Algorithm for High-Performance Wave Modeling
- Scientific Supercomputing with Graphics Processing Units
- Resource management in a multicore operating system
- A Roofline Visualization Framework
- Predicting and Optimizing System Utilization and Performance via Statistical Machine Learning
- Optimization Techniques for Mapping Algorithms and Applications onto CUDA GPU Platforms and CPU-GPU Heterogeneous Platforms
- Accelerating low-fidelity aerodynamic codes on multi- and many-core architectures
- Parallel Application Library for Object Recognition
Related papers
- Ikonic imagery in the severely subnormal.
- Study on Matrix Transpose Algorithms
- Research and implementation of matrix transpose for real-time SAR imaging system
- Adaptive Modified Transpose Jacobian Control of a Two-fingered Robotic Hand
- A note on positive partial transpose blocks
- Accessory cusp: Cusp of carabelli - A brief review
- Entangled States with Positive Partial Transpose in Any-Dimensional Quantum System