Open AccessBook
Scalable Analytical Query Processing
Martina-Cezara Albutiu
- 21 Dec 2013
2
TL;DR: This work proposes to avoid possibly wrong optimizer decisions regarding physical join operators by replacing these operators by a single one, called g-join, and develops a suite of massively parallel sort-merge (MPSM) join algorithms, which exploit the parallelization potential of multi-core CPUs.
read more
Abstract: Analytical query processing in database management systems aims at providing information within an acceptable time while affecting the performance of concurrent transactional workloads as little as possible. Scalability refers to the ability of database management systems to take advantage of additional resources to improve performance. In order to achieve the goal of scalable analytical query processing, we examine three approaches: First, we focus on synergy-based workload management, which exploits synergies between concurrently executed queries in order to maximize performance. Second, we concentrate on the robust execution of single queries. We propose to avoid possibly wrong optimizer decisions regarding physical join operators by replacing these operators by a single one, called g-join. Third, we turn our attention toward modern architectures and develop a suite of massively parallel sort-merge (MPSM) join algorithms, which exploit the parallelization potential of multi-core CPUs.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Have query optimizers hit the wall
TL;DR: Evidence is presented for a previously unknown upper bound on the number of operators a DBMS may be able to support before performance suffers, and it is shown that this upper bound may have already been reached.
5
Cache-Efficient Aggregation: Hashing Is Sorting
Ingo Müller,Peter Sanders,Arnaud Lacurie,Wolfgang Lehner,Franz Färber +4 more
- 27 May 2015
TL;DR: This paper argues that in terms of cache efficiency, the two paradigms of hashing and sorting are actually the same, and designs an algorithmic framework that allows to switch seamlessly between hashing and sorted routines during execution.
References
MapReduce: simplified data processing on large clusters
Jeffrey Dean,Sanjay Ghemawat +1 more
- 06 Dec 2004
TL;DR: This paper presents the implementation of MapReduce, a programming model and an associated implementation for processing and generating large data sets that runs on a large cluster of commodity machines and is highly scalable.
Query evaluation techniques for large databases
TL;DR: This survey describes a wide array of practical query evaluation techniques for both relational and postrelational database systems, including iterative execution of complex query evaluation plans, the duality of sort- and hash-based set-matching algorithms, types of parallel query execution and their implementation, and special operators for emerging database application domains.
Implementation techniques for main memory database systems
David J. DeWitt,Randy H. Katz,Frank Olken,Leonard D. Shapiro,Michael Stonebraker,Darien Wood +5 more
- 01 Jun 1984
TL;DR: This paper considers the changes necessary to permit a relational database system to take advantage of large amounts of main memory, and evaluates AVL vs B+-tree access methods, hash-based query processing strategies vs sort-merge, and study recovery issues when most or all of the database fits in main memory.
HyPer: A hybrid OLTP&OLAP main memory database system based on virtual memory snapshots
Alfons Kemper,Thomas Neumann +1 more
- 11 Apr 2011
TL;DR: This work presents an efficient hybrid system, called HyPer, that can handle both OLTP and OLAP simultaneously by using hardware-assisted replication mechanisms to maintain consistent snapshots of the transactional data.
755
SAP HANA database: data management for modern business applications
Franz Färber,Sang Kyun Cha,Jürgen Primsch,Christof Bornhövd,Stefan Sigg,Wolfgang Lehner +5 more
- 11 Jan 2012
TL;DR: The SAP HANA database permits the exchange of application semantics with the underlying data management platform that can be exploited to increase query expressiveness and to reduce the number of individual application-to-database round trips.
531
Related Papers (5)
Ram Gopal
- 11 Jan 1993
Sailesh Krishnamurthy,Chung Wu,Michael J. Franklin +2 more
- 27 Jun 2006
Miao Jiang,Ye Wang +1 more
- 13 Dec 2010