AutoFDO: automatic feedback-directed optimization for warehouse-scale applications
Dehao Chen,David Xinliang Li,Tipp Moseley +2 more
- 29 Feb 2016
- pp 12-23
TL;DR: AutoFDO is a system to simplify real-world deployment of feedback-directed optimization (FDO) by sampling hardware performance monitors on production machines and using those profiles to guide optimization.
read more
Abstract: AutoFDO is a system to simplify real-world deployment of feedback-directed optimization (FDO). The system works by sampling hardware performance monitors on production machines and using those profiles in to guide optimization. Profile data is stale by design, and we have implemented compiler features to deliver stable speedup across releases. The resulting performance is a geomean of 10.5% improvement on our benchmarks. AutoFDO achieves 85% of the gains of traditional FDO, despite imprecision due to sampling and information lost in the compilation pipeline. The system is deployed to hundreds of binaries at Google, and it is extremely easy to enable; users need only to add some flags to their release build. To date, AutoFDO has increased the number of FDO users at Google by 8X and has doubled the number of cycles spent in FDO-optimized binaries. Over half of CPU cycles used are now spent in some flavor of FDO-optimized binaries.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
VESPA: static profiling for binary optimization
Angélica Aparecida Moreira,Guilherme Ottoni,Fernando Magno Quintão Pereira +2 more
- 15 Oct 2021
TL;DR: In this paper, the authors revisited the static profiling technique proposed by Calder et al. in the late 90s, and investigated its application to drive binary optimizations, in the context of the BOLT binary optimizer, as a replacement for dynamic profiling.
MLOS: An Infrastructure for Automated Software Performance Engineering
Carlo Curino,Neha Godwal,Brian Kroth,Sergiy Kuryata,Greg Lapinski,Siqi Liu,Slava Oks,Olga Poppe,Adam Smiechowski,Ed Thayer,Markus Weimer,Yiwen Zhu +11 more
TL;DR: MLOS is a Data Science powered infrastructure and methodology to democratize and automate Software Performance Engineering that enables continuous, instance-level, robust, and trackable systems optimization.
Optimistic and Scalable Global Function Merging
Kyungwoo Lee,Manman Ren,Ellis Hoag +2 more
- 20 Jun 2024
TL;DR: Global function merging optimizes code size and build time by merging identical functions, but traditional approaches require complete intermediate representation. This paper introduces a scalable global function merger leveraging global merge information to create merging instances independently within each module context.
References
MapReduce: simplified data processing on large clusters
Jeffrey Dean,Sanjay Ghemawat +1 more
- 06 Dec 2004
TL;DR: This paper presents the implementation of MapReduce, a programming model and an associated implementation for processing and generating large data sets that runs on a large cluster of commodity machines and is highly scalable.
MapReduce: simplified data processing on large clusters
Jeffrey Dean,Sanjay Ghemawat +1 more
TL;DR: This presentation explains how the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks.
Bigtable: a distributed storage system for structured data
Fay W. Chang,Jeffrey Dean,Sanjay Ghemawat,Wilson C. Hsieh,Deborah A. Wallach,Michael Burrows,Tushar Deepak Chandra,Andrew Fikes,Robert E. Gruber +8 more
- 06 Nov 2006
TL;DR: Bigtable as discussed by the authors is a distributed storage system for managing structured data that is designed to scale to a very large size: petabytes of data across thousands of commodity servers, including web indexing, Google Earth and Google Finance.
Large-scale cluster management at Google with Borg
Abhishek Verma,Luis Pedrosa,Madhukar R. Korupolu,David Oppenheimer,Eric S. Tune,John Wilkes +5 more
- 17 Apr 2015
TL;DR: A summary of the Borg system architecture and features, important design decisions, a quantitative analysis of some of its policy decisions, and a qualitative examination of lessons learned from a decade of operational experience with it are presented.
Dremel: interactive analysis of web-scale datasets
Sergey Melnik,Andrey Gubarev,Jing Jing Long,Geoffrey M. Romer,Shiva Shivakumar,Matthew B. Tolton,Theodore Vassilakis +6 more
- 01 Sep 2010
TL;DR: The architecture and implementation of Dremel are described, and how it complements MapReduce-based computing is explained, and a novel columnar storage representation for nested records is presented.
Related Papers (5)
Chris Lattner,Vikram Adve +1 more
- 20 Mar 2004
Grant Ayers,Jung Ho Ahn,Christos Kozyrakis,Parthasarathy Ranganathan +3 more
- 01 Feb 2018