Automatic Sublining for Efficient Sparse Memory Accesses
TL;DR: The Instruction Spatial Locality Estimator (ISLE) as discussed by the authors is a hardware detector that finds instructions that access isolated words in a sea of unused data, while keeping regular accesses cached.
read more
Abstract: Sparse memory accesses, which are scattered accesses to single elements of a large data structure, are a challenge for current processor architectures. Their lack of spatial and temporal locality and their irregularity makes caches and traditional stream prefetchers useless. Furthermore, performing standard caching and prefetching on sparse accesses wastes precious memory bandwidth and thrashes caches, deteriorating performance for regular accesses. Bypassing prefetchers and caches for sparse accesses, and fetching only a single element (e.g., 8 B) from main memory (subline access), can solve these issues. Deciding which accesses to handle as sparse accesses and which as regular cached accesses, is a challenging task, with a large potential impact on performance. Not only is performance reduced by treating sparse accesses as regular accesses, not caching accesses that do have locality also negatively impacts performance by significantly increasing their latency and bandwidth consumption. Furthermore, this decision depends on the dynamic environment, such as input set characteristics and system load, making a static decision by the programmer or compiler suboptimal. We propose the Instruction Spatial Locality Estimator (ISLE), a hardware detector that finds instructions that access isolated words in a sea of unused data. These sparse accesses are dynamically converted into uncached subline accesses, while keeping regular accesses cached. ISLE does not require modifying source code or binaries, and adapts automatically to a changing environment (input data, available bandwidth, etc.). We apply ISLE to a graph analytics processor running sparse graph workloads, and show that ISLE outperforms the performance of no subline accesses, manual sublining, and prior work on detecting sparse accesses.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
The Intel® Programmable and Integrated Unified Memory Architecture (PIUMA) Graph Analytics Processor
TL;DR: The PIUMA architecture is presented and the experience in designing and building a prototype chip and its bring-up process are documents, some of which will be incorporated into future Intel products.
5
McCore: A Holistic Management of High-Performance Heterogeneous Multicores
Jaewon Kwon,Yongju Lee,Hongju Kal,Minjae Kim,Youngsok Kim,Won Woo Ro +5 more
- 28 Oct 2023
TL;DR: The McCore structure aims to enhance performance by partitioning the shared LLC based on each cluster’s asymmetric computing power and conditionally enabling fine-grained access in small cores, and significantly outperforms existing state-of-the-art cache partitioning, sparse access managing schemes, and heterogeneous multicore schedulers.
The First Direct Mesh-to-Mesh Photonic Fabric
Jason D. Howard,Joshua B. Fryman,Shamsul Abedin +2 more
TL;DR: The first direct mesh-to-mesh photonic fabric enables high-bandwidth sparse graph analytics by integrating optical transceivers directly into the compute logic.
References
Understanding DDR4 in pursuit of In-DRAM ECC
Sanghyuk Kwon,Young Hoon Son,Jung Ho Ahn +2 more
- 01 Nov 2014
TL;DR: The possibilities and challenges of implementing In-DRAM ECC on DDR4 SDRAM devices are identified and an ALERT_n pad is introduced, which can be used to report errors detected by a SECDED code in DRAM.
25
Dynamic fine-grained sparse memory accesses
Berkin Akin,Chiachen Chou,Jongsoo Park,Christopher J. Hughes,Rajat Agarwal +4 more
- 01 Oct 2018
TL;DR: This work proposes an ISA and microarchitecture that dynamically adapts to the locality of an input, and selectively exploits the fine-grained memory access capability provided by recently introduced memory technologies such as HBM and HMC.
3
Adaptive GPU cache bypassing
Yingying Tian,Sooraj Puthoor,Joseph L. Greathouse,Bradford M. Beckmann,Daniel A. Jimenez +4 more
- 07 Feb 2015
TL;DR: A GPU cache management technique that adaptively bypasses the GPU cache for blocks that are unlikely to be referenced again before being evicted, resulting in better performance for programs that do not use programmer-managed scratchpad memories.
Adaptive granularity memory systems: a tradeoff between storage efficiency and throughput
Doe Hyun Yoon,Min Kyu Jeong,Mattan Erez +2 more
- 04 Jun 2011
TL;DR: The evaluation shows that performance is improved by 61% without ECC and 44% with ECC in memory-intensive applications, while the reduction in memory power consumption and traffic is significant.
Locality-Driven Dynamic GPU Cache Bypassing
Chao Li,Shuaiwen Leon Song,Hongwen Dai,Albert Sidelnik,Siva Kumar Sastry Hari,Huiyang Zhou +5 more
- 08 Jun 2015
TL;DR: This paper presents a design that integrates locality filtering based on reuse characteristics of GPU workloads into the decoupled tag store of the existing L1 D-cache through simple and cost-effective hardware extensions.
Related Papers (5)
Xiangyao Yu,Christopher J. Hughes,Nadathur Satish,Srinivas Devadas +3 more
- 05 Dec 2015
Sam Ainsworth,Timothy M. Jones +1 more
- 01 Jun 2016
Liuxi Yang,Josep Torrellas +1 more
- 01 Feb 1997