Distributed data cache designs for clustered VLIW processors
TL;DR: This paper proposes and evaluates three different configurations of the L1 data cache among clusters for clustered VLIW processors, a snoop-based cache coherence scheme, a word-interleaved cache, and flexible LO-buffers managed by the compiler, and shows that the performance of such fully distributed architectures is always better than theperformance of a partially distributed one with the same amount of resources.
read more
Abstract: Wire delays are a major concern for current and forthcoming processors. One approach to deal with this problem is to divide the processor into semi-independent units referred to as clusters. A cluster usually consists of a local register file and a subset of the functional units, while the L1 data cache typically remains centralized in What we call partially distributed architectures. However, as technology evolves, the relative latency of such a centralized cache will increase, leading to an important impact on performance. In this paper, we propose partitioning the L1 data cache among clusters for clustered VLIW processors. We refer to this kind of design as fully distributed processors. In particular; we propose and evaluate three different configurations: a snoop-based cache coherence scheme, a word-interleaved cache, and flexible LO-buffers managed by the compiler. For each alternative, instruction scheduling techniques targeted to cyclic code are developed. Results for the Mediabench suite'show that the performance of such fully distributed architectures is always better than the performance of a partially distributed one with the same amount of resources. In addition, the key aspects of each fully distributed configuration are explored.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Patent
Specifying an access hint for prefetching partial cache block data in a cache hierarchy
Bradly G. Frey,Guy Lynn Guthrie,Cathy May,Ramakrishnan Rajamony,Balaram Sinharoy,William J. Starke,Peter K. Szwed +6 more
- 16 Apr 2009
TL;DR: In this paper, the authors propose a system and method for specifying an access hint for prefetching only a subsection of cache block data, for more efficient system interconnect usage by the processor core.
23
Patent
Updating Partial Cache Lines in a Data Processing System
David W. Cummings,Guy Lynn Guthrie,Hugh Shen,William J. Starke,Derek Edward Williams,Phillip G. Williams +5 more
- 15 Apr 2009
TL;DR: In this article, the authors propose a multi-level cache hierarchy coupled to and supporting the processor core, which includes at least one upper level of cache memory having a lower access latency and a higher access latency.
16
Inter-cluster communication in VLIW architectures
Andrei Terechko,Henk Corporaal +1 more
TL;DR: It is revealed that achieving the best characteristics of a clustered VLIW requires a thorough selection of an Inter-cluster Communication (ICC) model, which is the way clustering is exposed in the Instruction Set Architecture.
16
Clustered VLIW architectures : a quantitative approach
Andrei Terechko
- 01 Jan 2007
TL;DR: The final author version and the galley proof are versions of the publication after peer review that features the final layout of the paper including the volume, issue and page numbers.
Patent
Dynamic selection of a memory access size
Lakshminarayana B. Arimilli,Ravi Kumar Arimilli,Jerry Don Lewis,Warren E. Maule +3 more
- 01 Feb 2008
TL;DR: In this paper, the processing unit monitors utilization of data accessed by the plurality of memory accesses, and dynamically alters a memory access mode of operation so that a subsequent storage-modifying memory access targets less than a full cache line of data.
11
References
Complexity-effective superscalar processors
Subbarao Palacharla,Norman P. Jouppi,James E. Smith +2 more
- 01 May 1997
TL;DR: A microarchitecture that simplifies wakeup and selection logic is proposed and discussed, which will help minimize performance degradation due to slow bypasses in future wide-issue machines.
Iterative module scheduling: an algorithm for software pipelining loops
B. Ramakrishna Rau
- 30 Nov 1994
TL;DR: This paper presents a practical algorithm, iterative modulo scheduling, that is capable of dealing with realistic machine models and characterizes the algorithm in terms of the quality of the generated schedules as well the computational expense incurred.
749
Baring it all to software: Raw machines
E. Waingold,Michael Taylor,Devabhaktuni Srikrishna,Vivek Sarkar,Whay S. Lee,Victor W. Lee,Jason Kim,Matthew I. Frank,P. Finch,Rajeev Barua,Jonathan Babb,Saman Amarasinghe,Anant Agarwal +12 more
TL;DR: The most radical of the architectures that appear in this issue are Raw processors-highly parallel architectures with hundreds of very simple processors coupled to a small portion of the on-chip memory, allowing synthesis of complex operations directly in configured hardware.
725
The microarchitecture of the Pentium 4 processor
G. Hinton
- 01 Jan 2001
TL;DR: The main features and functions of the NetBurst microarchitecture of Intel’s new flagship Pentium 4 processor are described, including its new form of instruction cache called the Execution Trace Cache.
671
Effective compiler support for predicated execution using the hyperblock
Scott Mahlke,David C. Lin,William Y. Chen,Richard E. Hank,Roger A. Bringmann +4 more
- 10 Dec 1992
TL;DR: In this paper, a new structure, referred to as the hyperblock, is proposed to combine speculative execution with predicated execution for both compile-time optimization and scheduling of conditional branches.