On Optimally Partitioning Variable-Byte Codes
TL;DR: In this paper, the authors show that the compression ratio of Variable-Byte can be improved by adopting a partitioned representation of the inverted lists, which makes variable-Byte surprisingly competitive in space with the best bit-aligned encoders, hence disproving the folklore belief that variable-byte is space-inefficient for inverted index compression.
read more
Abstract: The ubiquitous Variable-Byte encoding is one of the fastest compressed representation for integer sequences. However, its compression ratio is usually not competitive with other more sophisticated encoders, especially when the integers to be compressed are small which is the typical case for inverted indexes. This paper shows that the compression ratio of Variable-Byte can be improved by $2\times$ 2 × by adopting a partitioned representation of the inverted lists. This makes Variable-Byte surprisingly competitive in space with the best bit-aligned encoders, hence disproving the folklore belief that Variable-Byte is space-inefficient for inverted index compression. Despite the significant space savings, we show that our optimization almost comes for free, given that: we introduce an optimal partitioning algorithm that does not affect indexing time because of its linear-time complexity; we show that the query processing speed of Variable-Byte is preserved, with an extensive experimental analysis and comparison with several other state-of-the-art encoders.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Techniques for Inverted Index Compression
TL;DR: The aim of this article is surveying the encoding algorithms suitable for inverted index compression and characterizing the performance of the inverted index through experimentation.
77
CC-News-En: A Large English News Corpus
Joel Mackenzie,Rodger Benham,Matthias Petri,Johanne R. Trippas,J. Shane Culpepper,Alistair Moffat +5 more
- 19 Oct 2020
TL;DR: A static, open-access news corpus is described using data from the Common Crawl Foundation, who provide free, publicly available web archives, including a continuous crawl of international news articles published in multiple languages, to support offline effectiveness experiments and hence batch evaluation campaigns.
64
On weighted k-mer dictionaries
TL;DR: In this paper , a weighted dictionary of [Formula: see text]-mers and their abundance counts, or weights, is proposed to reduce the number of runs in the weights to improve compression.
•Posted Content
Techniques for Inverted Index Compression
TL;DR: In this article, the authors survey the encoding algorithms suitable for inverted index compression and characterize the performance of the inverted index through experimentation, showing that index compression is essential because it leads to better exploitation of the computer memory hierarchy for faster query processing and, at the same time, allows reducing the number of storage machines.
7
•Posted Content
On Slicing Sorted Integer Sequences
TL;DR: A solution is proposed and implemented that recursively slices the universe of representation of a sequence to achieve compact storage and attain to fast query execution, thus offering an excellent space/time trade-off for the problem.
References
Inverted index compression and query processing with optimized document ordering
Hao Yan,Shuai Ding,Torsten Suel +2 more
- 20 Apr 2009
TL;DR: This work performs an extensive study of compression techniques for document IDs and presents new optimizations of existing techniques which can achieve significant improvement in both compression and decompression performances.
Efficient Storage and Retrieval by Content and Address of Static Files
TL;DR: Firm lower bounds are given to minimax measures of bits stored and bits accessed for each of four retrieval questions, and representations and algorithms for a bit-addressable machine which come within factors of two or three of attaining all four bounds at once for files of any size.
283
Compressing Integers for Fast File Access
Hugh E. Williams,Justin Zobel +1 more
TL;DR: It is shown experimentally that, for large or small collections, storing integers in a compressed format reduces the time required for either sequential stream access or random access.
267
The Design and Implementation of Modern Column-Oriented Database Systems
TL;DR: The design and implementation of modern column-oriented database systems can be found in this paper, with a specific focus on three influential research prototypes, MonetDB, C-Store, and X100, which form the basis for several well-known commercial column-store implementations.
SkimpyStash: RAM space skimpy key-value store on flash-based storage
Biplob Debnath,Sudipta Sengupta,Jin Li +2 more
- 12 Jun 2011
TL;DR: SkimpyStash as mentioned in this paper uses a hash table directory in RAM to index key-value pairs stored in a log-structured manner on flash, and moves most of the pointers that locate each keyvalue pair from RAM to flash itself.
Related Papers (5)
Peter Chovanec,Michal Kratky,Jiri Walder +2 more
- 30 Dec 2010
Alistair Moffat,JS Culpepper +1 more
- 01 Jan 2007
Fabian J. Corrales,David Chiu,Jason Sawin +2 more
- 29 Aug 2011