Book Chapter10.1007/978-3-319-17473-0_23
Exploring and Evaluating Array Layout Restructuring for SIMDization
Christopher Haine,Olivier Aumage,Enguerrand Petit,Denis Barthou +3 more
- 15 Sep 2014
- pp 351-366
TL;DR: This paper explains why modern compilers may fail or SIMDize poorly, due to conservativeness, source complexity or missing capabilities, in relation to SIMD processor units.
read more
Abstract: SIMD processor units have become ubiquitous. Using SIMD instructions is the key for performance for many applications. Modern compilers have made immense progress in generating efficient SIMD code. However, they still may fail or SIMDize poorly, due to conservativeness, source complexity or missing capabilities. When SIMDization fails, programmers are left with little clues about the root causes and actions to be taken.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
CERE: LLVM-Based Codelet Extractor and REplayer for Piecewise Benchmarking and Optimization
TL;DR: This article presents Codelet Extractor and REplayer (CERE), an open-source framework for code isolation that finds and extracts the hotspots of an application as isolated fragments of code, called codelets, and uses them in a realistic study to evaluate three different architectures on the NAS benchmarks.
•Proceedings Article
Tools for High Performance Computing: Proceedings of the 2nd International Workshop on Parallel Tools for High Performance Computing, July 2008, HLRS, Stuttgart
Michael Resch,Rainer Keller,Valentin Himmler,Bettina Krammer,Alexander Schulz +4 more
- 06 Aug 2008
TL;DR: The 2nd Parallel Tools Workshop as mentioned in this paper provides an overview of the existing tools in the area of integrated development environments for clusters, various parallel debuggers, and new-style performance analysis tools, as well as an update on the state of the art of long-term research tools.
31
Middleware Power Saving Scheme for Mobile Applications
Noor Zaman,Fatimah Abdualaziz Almusalli,Sarfarz N Brohi,Azween Abdullah +3 more
- 01 Oct 2018
TL;DR: This research introduces memory optimization through middleware transformation service, which converts the AOS to SOA and resulting will increase significant power saving for mobile application by reducing the memory access counts, which extended the battery life by minimizing its use.
13
Energy efficient middleware: Design and development for mobile applications
Fatimah Abdualaziz Almusalli,Noor Zaman,Raihan Ur Rasool +2 more
- 01 Jan 2017
TL;DR: The service will convert the data layout in memory from AOS to SOA, which will reduce the power consumed by memory and processor and result in efficient and extended battery life.
11
•Dissertation
Combiner approches statique et dynamique pour modéliser la performance de boucles HPC
Vincent Palomares
- 21 Sep 2015
TL;DR: UFS is described, an approach combining static analysis and cycle accurate simulation to very quickly estimate a loop’s execution time while accounting for out-of-order limitations in modern CPUs.
4
References
Optimizing data permutations for SIMD devices
Gang Ren,Peng Wu,David Padua +2 more
- 11 Jun 2006
TL;DR: A strategy to optimize all forms of data permutations is presented and it is shown that up to 77% of the permutation instructions are eliminated and, as a result, the average performance improvement is 48% on VMX and 68% on SSE2.
102
Tools for High Performance Computing
Michael Resch,Rainer Keller,Valentin Himmler,Bettina Krammer,Alexander Schulz +4 more
- 01 Jan 2008
TL;DR: This workshop will give the users an overview of the existing tools in the area of integrated development environments for clusters, various parallel debuggers, and new-style performance analysis tools, as well as an update on the state of the art of long-term research tools, which have advanced to an industrial level.
94
•Book
Tools for High Performance Computing 2009
Matthias S. Müller,Michael Resch,Alexander Schulz,Wolfgang E. Nagel +3 more
- 01 Jan 2010
62
Prediction and trace compression of data access addresses through nested loop recognition
Alain Ketterlin,Philippe Clauss +1 more
- 06 Apr 2008
TL;DR: An algorithm that takes a trace as input, and from that produces a sequence of loop nests that, when run, produces exactly the original sequence that is suitable for any kind of program execution trace.
Relaxing SIMD control flow constraints using loop transformations
Reinhard von Hanxleden,Ken Kennedy +1 more
- 01 Jul 1992
TL;DR: It is argued that loop flattening, whether performed by the programmer or by the compiler, introduces negligible overhead and can significantly improve the performance of scientific codes for solving irregular problems.
40