Hardware/compiler codevelopment for an embedded media processor
Christos Kozyrakis,David L. Judd,Joseph Gebis,Samuel Williams,David A. Patterson,Katherine Yelick +5 more
- 01 Nov 2001
- Vol. 89, Iss: 11, pp 1694-1709
TL;DR: This paper presents the codevelopment of the instruction set, the hardware, and the compiler for the Vector IRAM media processor, and describes how the architecture, design, and compiler features come together in a prototype system-on-a-chip, able to execute 3.2 billion operations per second per watt.
read more
Abstract: Embedded and portable systems running multimedia applications create a new challenge for hardware architects. A microprocessor for such applications needs to be easy to program like a general-purpose processor and have the performance and power efficiency of a digital signal processor. This paper presents the codevelopment of the instruction set, the hardware, and the compiler for the Vector IRAM media processor. A vector architecture is used to exploit the data parallelism of multimedia programs, which allows the use of highly modular hardware and enables implementations that combine high performance, low power consumption, and reduced design complexity. It also leads to a compiler model that is efficient both in terms of performance and executable code size. The memory system for the vector processor is implemented using embedded DRAM technology, which provides high bandwidth in an integrated, cost-effective manner. The hardware and the compiler for this architecture make complementary contributions to the efficiency of the overall system. This paper explores the interactions and tradeoffs between them, as well as the enhancements to a vector architecture necessary for multimedia processing. We also describe how the architecture, design, and compiler features come together in a prototype system-on-a-chip, able to execute 3.2 billion operations per second per watt.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Outer-loop vectorization: revisited for short SIMD architectures
Dorit Nuzman,Ayal Zaks +1 more
- 25 Oct 2008
TL;DR: This paper revisit the method of outer loop vectorization, paying special attention to properties of modern short SIMD architectures, and presents an optimization tapping such opportunities, capable of further boosting the performance obtained by outer-loop vectorization to achieve average speedup factors of 5.26 and 3.64.
194
Vector vs. superscalar and VLIW architectures for embedded multimedia benchmarks
Christoforos Kozyrakis,David A. Patterson +1 more
- 18 Nov 2002
TL;DR: This paper uses EEMBC, an industrial benchmark suite, to compare the VIRAM vector architecture to superscalar and VLIW processors for embedded multimedia applications and demonstrates that executable code for VirAM is up to 10 times smaller than V LIW code and comparable to x86 CISC code.
A Programmable Vision Chip Based on Multiple Levels of Parallel Processors
TL;DR: A novel programmable vision chip based on multiple levels of parallel processors that can satisfy flexibly the needs of different vision applications such as image pre-processing, complicated feature extraction and over 1000 fps high-speed image capture is proposed.
103
Scalable Vector Media-processors for Embedded Systems
Christoforos Kozyrakis,David A. Patterson +1 more
- 01 Jan 2002
TL;DR: It is argued that it is possible to design processors that deliver high performance, have low energy consumption, and are simple to implement and it is demonstrated that the vector instructions in VIRAM can capture the data-level parallelism in multimedia tasks and lead to smaller code size than RISC, CISC, and VLIW architectures.
FPGA implementation and performance evaluation of a high throughput crypto coprocessor
TL;DR: The FPGA implementation of FastCrypto, which extends a general-purpose processor with a crypto coprocessor for encrypting/decrypting data, is described and the trade-offs between Fastcrypto performance and design parameters are studied, including the number of stages per round, thenumber of parallel Advance Encryption Standard (AES) pipelines, and the size of the queues.
46
References
•Book
Computer Architecture: A Quantitative Approach
John L. Hennessy,David A. Patterson +1 more
- 01 Dec 1989
TL;DR: This best-selling title, considered for over a decade to be essential reading for every serious student and practitioner of computer design, has been updated throughout to address the most important trends facing computer designers today.
12.6K
Low-power CMOS digital design
TL;DR: In this paper, techniques for low power operation are presented which use the lowest possible supply voltage coupled with architectural, logic style, circuit, and technology optimizations to reduce power consumption in CMOS digital circuits while maintaining computational throughput.
The future of wires
R. Ho,Ken Mai,Mark Horowitz +2 more
- 01 Apr 2001
TL;DR: Wires that shorten in length as technologies scale have delays that either track gate delays or grow slowly relative to gate delays, which is good news since these "local" wires dominate chip wiring.
Compiler transformations for high-performance computing
TL;DR: This survey is a comprehensive overview of the important high-level program restructuring techniques for imperative languages, such as C and Fortran, and describes the purpose of each transformation, how to determine if it is legal, and an example of its application.
1K
Related Papers (5)
John L. Hennessy,David A. Patterson +1 more
- 01 Dec 1989
K. Diefendorff,Pradeep Dubey +1 more
James E. Smith,Gurindar S. Sohi +1 more
- 01 Dec 1995
[...]
R. Ho,Ken Mai,Mark Horowitz +2 more
- 01 Apr 2001