A high-performance algorithm for identifying frequent items in data streams
Daniel Anderson,Pryce Bevan,Kevin J. Lang,Edo Liberty,Lee Rhodes,Justin Thaler +5 more
- 01 Nov 2017
- pp 268-282
39
TL;DR: A highly optimized version of Misra and Gries' algorithm for estimating frequencies of items over data streams that is suitable for deployment in industrial settings and improves on two theoretical and practical aspects of prior work.
read more
Abstract: Estimating frequencies of items over data streams is a common building block in streaming data measurement and analysis. Misra and Gries introduced their seminal algorithm for the problem in 1982, and the problem has since been revisited many times due its practicality and applicability. We describe a highly optimized version of Misra and Gries' algorithm that is suitable for deployment in industrial settings. Our code is made public via an open source library called Data Sketches that is already used by several companies and production systems. Our algorithm improves on two theoretical and practical aspects of prior work. First, it handles weighted updates in amortized constant time, a common requirement in practice. Second, it uses a simple and fast method for merging summaries that asymptotically improves on prior work even for unweighted streams. We describe experiments confirming that our algorithms are more efficient than prior proposals.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Network-wide routing-oblivious heavy hitters
Ran Ben Basat,Gil Einziger,Shir Landau Feibish,Jalil Moraney,Danny Raz +4 more
- 23 Jul 2018
TL;DR: This work suggests the first network-wide and routing oblivious algorithms for three fundamental network monitoring problems, and provides a general, constant time framework that solves the distributed versions of the volume estimation, frequency estimation and heavy-hitters problems with provable guarantees.
41
Memento: making sliding windows efficient for heavy hitters
Ran Ben Basat,Gil Einziger,Isaac Keslassy,Ariel Orda,Shay Vargaftik,Erez Waisbard +5 more
- 04 Dec 2018
TL;DR: In this paper, the authors proposed a sliding window based algorithm for detecting heavy hitters through sliding windows, which achieved up to 273X faster detection than existing window-based techniques. But, existing techniques are slow in detecting new heavy hitters.
38
Secure Multi-party Computation of Differentially Private Heavy Hitters
Jonas Böhler,Florian Kerschbaum +1 more
- 12 Nov 2021
TL;DR: In this paper, the authors present a differentially private top-k algorithm for very small data sets (hundreds of values) using semi-honest computation parties distributed over the Internet.
38
Randomized Admission Policy for Efficient Top-k, Frequency, and Volume Estimation
TL;DR: This paper introduces Randomized Admission Policy –a novel algorithm for the frequency, top-k, and byte volume estimation problems, which are fundamental in network monitoring.
33
Stream frequency over interval queries
Ran Ben Basat,Roy Friedman,Rana Shahout +2 more
- 01 Dec 2018
TL;DR: This paper considers a generalized sliding window model that supports stream frequency queries over an interval given at query time, and asymptotically improves the space bounds of existing work, reduces the update and query time to a constant, and provides deterministic solutions.
References
•Book
Probability and Computing: Randomized Algorithms and Probabilistic Analysis
Michael Mitzenmacher,Eli Upfal +1 more
- 01 Jan 2005
TL;DR: Preface 1. Events and probability 2. Discrete random variables and expectation 3. Moments and deviations 4. Chernoff bounds 5. Balls, bins and random graphs 6. Probabilistic method 7. Markov chains and random walks 8. Continuous distributions and the Poisson process
2.7K
An improved data stream summary: the count-min sketch and its applications
Graham Cormode,S. Muthukrishnan +1 more
TL;DR: In this paper, the authors introduce a sublinear space data structure called the countmin sketch for summarizing data streams, which allows fundamental queries in data stream summarization such as point, range, and inner product queries to be approximately answered very quickly; in addition it can be applied to solve several important problems in data streams such as finding quantiles, frequent items, etc.
2.2K
Finding Frequent Items in Data Streams
Moses Charikar,Kevin Chen,Martin Farach-Colton +2 more
- 08 Jul 2002
TL;DR: This work presents a 1-pass algorithm for estimating the most frequent items in a data stream using limited storage space, which achieves better space bounds than the previously known best algorithms for this problem for several natural distributions on the item frequencies.
An improved data stream summary: The count-min sketch and its applications
Graham Cormode,S. Muthukrishnan +1 more
- 05 Apr 2004
TL;DR: The Count-Min Sketch allows fundamental queries in data stream summarization such as point, range, and inner product queries to be approximately answered very quickly and can be applied to solve several important problems in data streams such as finding quantiles, frequent items, etc.
Approximate frequency counts over data streams
Gurmeet Singh Manku,Rajeev Motwani +1 more
- 01 Aug 2012
TL;DR: This talk will trace the history of the Approximate Frequency Counts paper, how it was conceptualized and how it influenced data stream research.
1.4K