Efficiently mining long patterns from databases
Roberto J. Bayardo
- 01 Jun 1998
- Vol. 27, Iss: 2, pp 85-93
TL;DR: A pattern-mining algorithm that scales roughly linearly in the number of maximal patterns embedded in a database irrespective of the length of the longest pattern, compared with previous algorithms that scale exponentially with longest pattern length.
read more
Abstract: We present a pattern-mining algorithm that scales roughly linearly in the number of maximal patterns embedded in a database irrespective of the length of the longest pattern. In comparison, previous algorithms based on Apriori scale exponentially with longest pattern length. Experiments on real data show that when the patterns are long, our algorithm is more efficient by an order of magnitude or more.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Fast vertical mining using diffsets
Mohammed J. Zaki,Karam Gouda +1 more
- 24 Aug 2003
TL;DR: This paper presents a novel vertical data representation called Diffset, that only keeps track of differences in the tids of a candidate pattern from its generating frequent patterns, and shows that diffsets drastically cut down the size of memory required to store intermediate results.
A Tree Projection Algorithm for Generation of Frequent Item Sets
TL;DR: This paper provides an implementation of the tree projection method which is up to one order of magnitude faster than other recent techniques in the literature and has a well-structured data access pattern which provides data locality and reuse of data for multiple levels of the cache.
678
The RDF-3X engine for scalable management of RDF data
Thomas Neumann,Gerhard Weikum +1 more
- 01 Feb 2010
TL;DR: The RDF-3X engine is presented, an implementation of SPARQL that achieves excellent performance by pursuing a RISC-style architecture with streamlined indexing and query processing, and can outperform the previously best alternatives by one or two orders of magnitude.
CLOSET+: searching for the best strategies for mining frequent closed itemsets
Jianyong Wang,Jiawei Han,Jian Pei +2 more
- 24 Aug 2003
TL;DR: CLOSET+ integrates the advantages of the previously proposed effective strategies as well as some ones newly developed here, and develops a winning algorithm CLOSET+.
677
Automatic subspace clustering of high dimensional data for data mining applications
TL;DR: Data mining applications place special requirements on clustering algorithms including the ability to find clusters embedded in subspaces of high dimensional data, scalability, end-user comprehensiveness, and so on.
References
Mining association rules between sets of items in large databases
Rakesh Agrawal,Tomasz Imielinski,Arun N. Swami +2 more
- 01 Jun 1993
TL;DR: An efficient algorithm is presented that generates all significant association rules between items in the database of customer transactions and incorporates buffer management and novel estimation and pruning techniques.
•Proceedings Article
Fast algorithms for mining association rules
Rakesh Agrawal,Ramakrishnan Srikant +1 more
- 01 Jul 1998
TL;DR: Two new algorithms for solving thii problem that are fundamentally different from the known algorithms are presented and empirical evaluation shows that these algorithms outperform theknown algorithms by factors ranging from three for small problems to more than an order of magnitude for large problems.
Mining sequential patterns
Rakesh Agrawal,Ramakrishnan Srikant +1 more
- 06 Mar 1995
TL;DR: Three algorithms are presented to solve the problem of mining sequential patterns over databases of customer transactions, and empirically evaluating their performance using synthetic data shows that two of them have comparable performance.
Mining Sequential Patterns: Generalizations and Performance Improvements
Ramakrishnan Srikant,Ramakrishnan Srikant,Rakesh Agrawal +2 more
- 25 Mar 1996
TL;DR: This work adds time constraints that specify a minimum and/or maximum time period between adjacent elements in a pattern, and relax the restriction that the items in an element of a sequential pattern must come from the same transaction.
3.2K
•Proceedings Article
Fast discovery of association rules
Rakesh Agrawal,Heikki Mannila,Ramakrishnan Srikant,Hannu Toivonen,A. Inkeri Verkamo +4 more
- 01 Feb 1996
2.8K
Related Papers (5)
Jiawei Han,Jian Pei,Yiwen Yin +2 more
- 16 May 2000
Rakesh Agrawal,Ramakrishnan Srikant +1 more
- 01 Jul 1998
Rakesh Agrawal,Ramakrishnan Srikant +1 more
- 12 Sep 1994