Dynamic Support Range based Rare Pattern Mining over Data Streams
TL;DR: The Dynamic Support Range-based Hybrid-Eclat Algorithm (DSRHEA), an Eclat-based technique for mining unique patterns from a data stream using bit-set vertical mining with two item-based optimizations is developed.
read more
Abstract: —Rare itemset mining is a relatively recent topic of study in data mining. In certain application domains, such as online banking transaction analysis, sensor data analysis, and stock market analysis, rare patterns are patterns with low support and high confidence that are extremely interesting when compared to frequent patterns. Numerous applications generate large amounts of continuous data streams. We require efficient algorithms capable of processing data streams in order to analyze them and find unique patterns. The strategies developed for static databases cannot be used to data streams. As a result, we require algorithms created expressly for data stream processing in order to extract critical unique patterns. Rare pattern mining is still in its infancy, with only a few ways available. To address this is developed the Dynamic Support Range-based Hybrid-Eclat Algorithm (DSRHEA), an Eclat-based technique for mining unique patterns from a data stream using bit-set vertical mining with two item-based optimizations. The detected patterns are kept in a prefix-based rare pattern tree that uses double hashing to maintain the unusual pattern in the data stream. Testing showed that the proposed method did well in terms of how long it took to run, how many rare patterns it made and accuracy.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
References
A Method for Mining Infrequent Causal Associations and Its Application in Finding Adverse Drug Reaction Signal Pairs
TL;DR: An innovative data mining framework is proposed and a novel interestingness measure, exclusive causal-leverage, is created based on a computational, fuzzy recognition-primed decision (RPD) model that was previously developed to mine the causal relationship between drugs and their associated adverse drug reactions (ADRs).
Efficient discovery of approximate dependencies
Sebastian Kruse,Felix Naumann +1 more
- 01 Mar 2018
TL;DR: The novel and highly efficient algorithm P yro to discover both approximate FDs and approximate UCCs is proposed, which combines a separate-and-conquer search strategy with sampling-based guidance that quickly detects dependency candidates and verifies them.