Proceedings Article10.1109/HICSS.1998.648320
Exploiting parallelism in knowledge discovery systems to improve scalability
Gehad Galal,Diane J. Cook,Lawrence B. Holder +2 more
- 06 Jan 1998
- Vol. 5, pp 256-265
3
TL;DR: This research outlines a general approach for scaling KDD systems using parallel and distributed resources and applies the suggested strategies to the SUBDUE knowledge discovery system.
read more
Abstract: The large amount of data collected today is quickly overwhelming researchers' abilities to interpret the data and discover interesting patterns. Knowledge discovery and data mining approaches hold the potential to automate the interpretation process, but these approaches frequently utilize computationally expensive algorithms. In particular, scientific discovery systems focus on the utilization of richer data representation, sometimes without regard for scalability. This research outlines a general approach for scaling KDD systems using parallel and distributed resources and applies the suggested strategies to the SUBDUE knowledge discovery system. SUBDUE has been used to discover interesting and repetitive concepts in graph-based databases from a variety of domains, but requires a substantial amount of processing time. Experiments that demonstrate that scalability of parallel versions of the SUBDUE system are performed using CAD circuit databases and artificially-generated databases, and potential achievements and obstacles are discussed.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
User-centered system decomposition: Z-based requirements clustering
P. Hsia,C.T. Hsu,D.C. Kung,Lawrence B. Holder +3 more
- 15 Apr 1996
TL;DR: The paper presents a requirements clustering process based on ER modeling, scenarios, and the formal specification notation Z that produces a set of useful, usable, and semi-independent clusters that can be developed and delivered to the customers in increments.
10
Improvement of Query Processing Speed in Data Warehousing with the Usage of Components-Bitmap Indexing, Iceberg and Uncertain Data
Uma Pavan Kumar,Lakshma Reddy,Sreedevi. S. Erady +2 more
- 23 May 2015
TL;DR: The current article is dealing with the implementation of iceberg queries, slowly changing dimensions and uncertain data processing, which will improve the processing speed of the data warehousing.
Dynamics of modeling in data mining: interpretive approach to bankruptcy prediction
TL;DR: In this paper, the authors used a data-mining approach to develop bankruptcy prediction models suitable for normal and crisis economic conditions and observed the dynamics of model change from normal to crisis conditions and provided interpretation of bankruptcy classifications.
References
Knowledge acquisition via incremental conceptual clustering
TL;DR: COBWEB is a conceptual clustering system that organizes data so as to maximize inference ability, and is incremental and computationally economical, and thus can be flexibly applied in a variety of domains.
Multilevelk-way Partitioning Scheme for Irregular Graphs
George Karypis,Vipin Kumar +1 more
TL;DR: This paper presents and study a class of graph partitioning algorithms that reduces the size of the graph by collapsing vertices and edges, they find ak-way partitioning of the smaller graph, and then they uncoarsen and refine it to construct ak- way partitioning for the original graph.
2K
•Proceedings Article
Bayesian classification (AutoClass): theory and results
Peter Cheeseman,John Stutz +1 more
- 01 Feb 1996
TL;DR: It is emphasized that no current unsupervised classi cation system can produce maximally useful results when operated alone and that it is the interaction between domain experts and the machine searching over the model space that generates new knowledge.
1.2K
Knowledge discovery in databases
TL;DR: The past l0 years of KDD are described and predictions for the next 10 years are outlined and it is suggested that KDD should be renamed KDD2 or KDD3 to avoid confusion.
988