High Performance Stencil Code Algorithms for GPGPUs
Andreas Schäfer,Dietmar Fey +1 more
- 01 Jan 2011
- Vol. 4, pp 2027-2036
TL;DR: This paper represents the first successful application of temporal blocking for 3D stencils on GPGPUs and thereby exceeds previous results by a considerable margin and is also the first paper to study stencil codes on Fermi.
read more
Abstract: In this paper we investigate how stencil computations can be implemented on state-of-the-art general purpose graphics processing units (GPGPUs). Stencil codes can be found at the core of many numerical solvers and physical simulation codes and are therefore of particular interest to scientific computing research. GPGPUs have gained a lot of attention recently because of their superior floating point performance and memory bandwidth. Nevertheless, especially memory bound stencil codes have proven to be challenging for GPGPUs, yielding lower than to be expected speedups. We chose the Jacobi method as a standard benchmark to evaluate a set of algorithms on NVIDIA's latest Fermi chipset. One of our fastest algorithms is a parallel wavefront update. It exploits the enlarged on-chip shared memory to perform two time step updates per sweep. To the best of our knowledge, it represents the first successful applicationof temporal blocking for 3D stencils on GPGPUs and thereby exceeds previous results by a considerable margin. It is also the first paper to study stencil codes on Fermi.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
An analytical GPU performance model for 3D stencil computations from the angle of data traffic
TL;DR: A simple analytical model is proposed to estimate the execution time based on quantifying the data traffic volume at three stages: between registers and on-streaming multiprocessor (SMX) storage, between on-SMX storage and L2 cache, and between L2 caches and GPU’s device memory.
11
Evaluating optimizations that reduce global memory accesses of stencil computations in GPGPUs
TL;DR: Experimental data show that codes that use these optimizations are up to 3.3 times faster than the classical stencil formulation, and that the most profitable optimization varies with grid and stencil sizes.
11
•Posted Content
A Generic Library for Stencil Computations
Mauro Bianco,Ugo Varetto +1 more
TL;DR: A domain specific C++ generic library for stencil computations, like PDE solvers, which features high level constructs to specify computation and allows the development of parallel stencil Computations with very limited effort.
PIMS: a lightweight processing-in-memory accelerator for stencil computations
Jie Li,Xi Wang,Antonino Tumeo,Brody Williams,John D. Leidel,Yong Chen +5 more
- 30 Sep 2019
TL;DR: PIMS, implemented in the logic layer of a 3D-stacked memory, exploits the high bandwidth provided by through-silicon vias to reduce redundant memory traffic and is presented as an in-memory accelerator for stencil computations.
11
•Posted Content
MGSim + MGMark: A Framework for Multi-GPU System Research.
Yifan Sun,Trinayan Baruah,Saiful A. Mojumder,Shi Dong,Rafael Ubal,Xiang Gong,Shane Treadway,Yuhui Bao,Vincent Zhao,José L. Abellán,John Kim,Ajay Joshi,David Kaeli +12 more
TL;DR: MGSim, a cycle-accurate, extensively validated, multi-GPU simulator, based on AMD's Graphics Core Next 3 (GCN3) instruction set architecture is presented and the design implications from this work are evaluated, suggesting that D-MGPU is an attractive programming model for future multi- GPU systems.
References
•Book
Theory of Self-Reproducing Automata
John von Neumann,Arthur W. Burks +1 more
- 01 Jan 1966
TL;DR: This invention relates to prefabricated buildings and comprises a central unit having a peripheral section therearound to form a main residential part defined by an assembly of juxtaposed roofing and facing trusses.
5.7K
Theory of self-reproducing automata: John von Neumann (edited by A.W. Burks). University of Illinois Press, Urbana, 1966. xiii + 388pp., $10.00
TL;DR: This invention relates to prefabricated buildings and comprises a central unit having a peripheral section therearound to form a main residential part.
3.1K
Roofline: an insightful visual performance model for multicore architectures
TL;DR: The Roofline model offers insight on how to improve the performance of software and hardware in the rapidly changing world of connected devices.
3.5-D Blocking Optimization for Stencil Computations on Modern CPUs and GPUs
Anthony Nguyen,Nadathur Satish,Jatin Chhugani,Changkyu Kim,Pradeep Dubey +4 more
- 13 Nov 2010
TL;DR: A novel 3.
331