Book Chapter10.1007/978-3-319-06486-4_7
Intel Math Kernel Library
Endong Wang,Qing Zhang,Bo Shen,Guangyong Zhang,Xiaowei Lu,Qing Wu,Yajuan Wang +6 more
- 01 Jan 2014
- pp 167-188
730
TL;DR: In order to achieve optimal performance on multi-core and multi-processor systems, the features of parallelism and manage the memory hierarchical characters efficiently need to be used.
read more
Abstract: In order to achieve optimal performance on multi-core and multi-processor systems, we need to fully use the features of parallelism and manage the memory hierarchical characters efficiently. The performance of sequential codes relies on the instruction-level and register-level SIMD parallelism, and also on high-speed cache-blocking functions. Threading applications need advanced planning to achieve satisfactory load balancing.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Performance analysis of task-based multi-frontal sparse linear solvers: Structure matters
TL;DR: In this paper , the authors propose several performance visualization techniques and modeling strategies motivated by the analysis of task-based multifrontal sparse linear solvers whose structure is particularly complex, which can detect and highlight anomalies and understand resource utilization from the application point-of-view in a very insightful way.
1
On-the-Fly Lowering Engine: Offloading Data Layout Conversion for Convolutional Neural Networks
Mingu Kang,Sang Moo Hyun,Tae Hee Han,Jungrae Kim,Seokin Hong +4 more
TL;DR: A novel hardware mechanism, called OLE, which generates lowered matrix on-the-fly from the original input matrix to reduce memory footprint and bandwidth requirements and reduce the execution time of convolutional layers by 57.7% on average.
1
Third order tensor-oriented directional splitting for exponential integrators
Fabio Cassini,Marco Caliari +1 more
TL;DR: This work shows how to produce directional split approximations of third order with respect to the time step size and has been successfully tested against state-of-the-art techniques on two well-known physical models that lead to Turing patterns, namely the 2D Schnakenberg and the 3D FitzHugh--Nagumo systems.
Comparison of Reproducible Parallel Preconditioned BiCGSTAB Algorithm Based on ExBLAS and ReproBLAS
Xiao Min Lei,Tongxiang Gu,Stef Graillat,Xiaowen Xu,Jing Meng +4 more
- 27 Feb 2023
TL;DR: In this paper, the performance of the parallel preconditioned BiCGSTAB algorithm implemented with two different libraries (ExBLAS and ReproBLAS) that can ensure the reproducibility of computations is compared.
1
Optimizing CNNs on Multicores for Scalability, Performance and Goodput
TL;DR: Convolutional Neural Networks are a class of Ar- tificial Neural Networks that are highly efficient at the pattern recognition tasks that underlie difficult AI prob- lems in a variety of variety of situations.
1
References
Mersenne twister: a 623-dimensionally equidistributed uniform pseudo-random number generator
TL;DR: A new algorithm called Mersenne Twister (MT) is proposed for generating uniform pseudorandom numbers, which provides a super astronomical period of 2 and 623-dimensional equidistribution up to 32-bit accuracy, while using a working area of only 624 words.
•Book
The Art of Computer Programming, Volume 2: Seminumerical Algorithms
Donald E. Knuth
- 01 Jan 1981
4.4K
•Book
Non-uniform random variate generation
Luc Devroye
- 16 Apr 1986
TL;DR: A survey of the main methods in non-uniform random variate generation can be found in this article, where the authors provide information on the expected time complexity of various algorithms, before addressing modern topics such as indirectly specified distributions, random processes and Markov chain methods.
4K
Non-Uniform Random Variate Generation.
B. J. T. Morgan,Luc Devroye +1 more
TL;DR: This chapter reviews the main methods for generating random variables, vectors and processes in non-uniform random variate generation, and provides information on the expected time complexity of various algorithms before addressing modern topics such as indirectly specified distributions, random processes, and Markov chain methods.
3.7K
Related Papers (5)
Kaiming He,Xiangyu Zhang,Shaoqing Ren,Jian Sun +3 more
- 27 Jun 2016
Karen Simonyan,Andrew Zisserman +1 more
- 04 Sep 2014
Yousef Saad
- 01 Apr 2003