Proceedings Article10.1145/502512.502529
Mining time-changing data streams
Geoff Hulten,Laurie Spencer,Pedro Domingos +2 more
- 26 Aug 2001
- pp 97-106
TL;DR: An efficient algorithm for mining decision trees from continuously-changing data streams, based on the ultra-fast VFDT decision tree learner is proposed, called CVFDT, which stays current while making the most of old data by growing an alternative subtree whenever an old one becomes questionable, and replacing the old with the new when the new becomes more accurate.
read more
Abstract: Most statistical and machine-learning algorithms assume that the data is a random sample drawn from a stationary distribution. Unfortunately, most of the large databases available for mining today violate this assumption. They were gathered over months or years, and the underlying processes generating them changed during this time, sometimes radically. Although a number of algorithms have been proposed for learning time-changing concepts, they generally do not scale well to very large databases. In this paper we propose an efficient algorithm for mining decision trees from continuously-changing data streams, based on the ultra-fast VFDT decision tree learner. This algorithm, called CVFDT, stays current while making the most of old data by growing an alternative subtree whenever an old one becomes questionable, and replacing the old with the new when the new becomes more accurate. CVFDT learns a model which is similar in accuracy to the one that would be learned by reapplying VFDT to a moving window of examples every time a new example arrives, but with O(1) complexity per example, as opposed to O(w), where w is the size of the window. Experiments on a set of large time-changing data streams demonstrate the utility of this approach.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Predictive Handling of Asynchronous Concept Drifts in Distributed Environments
TL;DR: This work develops an ensemble approach, PINE, that combines reactive adaptation via drift detection, and proactive handling of upcoming changes via early warning and adaptation across the peers that handles asynchronous concept drifts better and faster than current state-of-the-art approaches.
•Proceedings Article
AIMS: An Immersidata Management System.
Cyrus Shahabi
- 01 Jan 2003
TL;DR: A system to address the challenges involved in managing the multidimensional sensor data streams generated within immersive environments by focusing on two applications, Attention De cit Hyperactivity Disorder (ADHD) diagnosis and American Sign Language (ASL) recognition is introduced.
Scarcity of Labels in Non-Stationary Data Streams: A Survey
TL;DR: The types of change, which can occur in a data-stream are formally described and the methods for dealing with change when there is limited access to labels are cataloged.
DSM-PLW: single-pass mining of path traversal patterns over streaming web click-sequences
TL;DR: This paper proposes a projection-based, single-pass algorithm for online incremental mining of path traversal patterns over a continuous stream of maximal forward references generated at a rapid rate, and has gently growing memory requirements and makes only one pass over the streaming data.
28
A survey of active and passive concept drift handling methods
TL;DR: Many concept drift handling methods in this survey are analyzed and summarized in terms of the comparing algorithms, learning model, applicable drift type, advantages, and disadvantages of the algorithms.
28
References
•Book
Classification and regression trees
Leo Breiman
- 01 Jan 1983
TL;DR: The methodology used to construct tree structured rules is the focus of a monograph as mentioned in this paper, covering the use of trees as a data analysis method, and in a more mathematical framework, proving some of their fundamental properties.
22.7K
On Estimation of a Probability Density Function and Mode
TL;DR: In this paper, the problem of the estimation of a probability density function and of determining the mode of the probability function is discussed. Only estimates which are consistent and asymptotically normal are constructed.
Probability Inequalities for sums of Bounded Random Variables
TL;DR: In this article, upper bounds for the probability that the sum S of n independent random variables exceeds its mean ES by a positive number nt are derived for certain sums of dependent random variables such as U statistics.
•Posted Content
On Information and Sufficiency
TL;DR: The information deviation between any two finite measures cannot be increased by any statistical operations (Markov morphisms) and is invarient if and only if the morphism is sufficient for these two measures as mentioned in this paper.
7.3K