A modified hyperplane clustering algorithm allows for efficient and accurate clustering of extremely large datasets
TL;DR: A new two-stage algorithm which partitions the high-dimensional space associated with microarray data using hyperplanes and reduces the memory requirements allowing us to cluster 44 460 genes without failure and significantly decreases the time to complete when compared with popular k-means programs.
read more
Abstract: Motivation: As the number of publically available microarray experiments increases, the ability to analyze extremely large datasets across multiple experiments becomes critical. There is a requirement to develop algorithms which are fast and can cluster extremely large datasets without affecting the cluster quality. Clustering is an unsupervised exploratory technique applied to microarray data to find similar data structures or expression patterns. Because of the high input/output costs involved and large distance matrices calculated, most of the algomerative clustering algorithms fail on large datasets (30 000 + genes/200 + arrays). In this article, we propose a new two-stage algorithm which partitions the high-dimensional space associated with microarray data using hyperplanes. The first stage is based on the Balanced Iterative Reducing and Clustering using Hierarchies algorithm with the second stage being a conventional k-means clustering technique. This algorithm has been implemented in a software tool (HPCluster) designed to cluster gene expression data. We compared the clustering results using the two-stage hyperplane algorithm with the conventional k-means algorithm from other available programs. Because, the first stage traverses the data in a single scan, the performance and speed increases substantially. The data reduction accomplished in the first stage of the algorithm reduces the memory requirements allowing us to cluster 44 460 genes without failure and significantly decreases the time to complete when compared with popular k-means programs. The software was written in C# (.NET 1.1).
Availability: The program is freely available and can be downloaded from http://www.amdcc.org/bioinformatics/bioinformatics.aspx.
Contact: [email protected]
Supplementary information:Supplementary data are available at Bioinformatics online.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Journal Article
Transcriptomic changes induced by mycophenolic acid in gastric cancer cells
TL;DR: It is suggested that MPA has beneficial anticancer activity through diverse molecular pathways and biological processes that include cell cycle, apoptosis, cell proliferation and migration.
28
Hepatic Gene Expression Profiling Reveals Key Pathways Involved in Leptin-Mediated Weight Loss in ob/ob Mice
Ashok Sharma,Shoshana M. Bartell,Clifton A. Baile,Bo Chen,Robert H. Podolsky,Richard A. McIndoe,Jin-Xiong She +6 more
TL;DR: Key molecular pathways and downstream genes which respond to leptin treatment and are involved in leptin-mediated weight loss are identified and further investigation will be required to assess the possible use of these genes and their associated protein products as therapeutic targets for the treatment of obesity.
Celda: a Bayesian model to perform co-clustering of genes into modules and cells into subpopulations using single-cell RNA-seq data
TL;DR: Celda as mentioned in this paper is a Bayesian hierarchical model to perform co-clustering of genes into transcriptional modules and cells into sub-populations for single-cell RNA-seq.
Celda: A Bayesian model to perform bi-clustering of genes into modules and cells into subpopulations using single-cell RNA-seq data
TL;DR: A novel Bayesian hierarchical model called Cellular Latent Dirichlet Allocation (Celda) is developed to perform bi-clustering of co-expressed genes into modules and cells into subpopulations and presents a novel principled approach towards characterizing transcriptional programs and cellular and heterogeneity in single-cell data.
CLIC: clustering analysis of large microarray datasets with individual dimension-based clustering
TL;DR: This study presents CLIC, which meets the requirements of clustering analysis particularly but not limited to large microarray data sets and enables iterative sub-clustering into more homogeneous groups and the identification of common expression patterns among the genes separated in different groups due to the large difference in the expression levels.
17
References
Cluster analysis and display of genome-wide expression patterns
TL;DR: A system of cluster analysis for genome-wide expression data from DNA microarray hybridization is described that uses standard statistical algorithms to arrange genes according to similarity in pattern of gene expression, finding in the budding yeast Saccharomyces cerevisiae that clustering gene expression data groups together efficiently genes of known similar function.
Objective Criteria for the Evaluation of Clustering Methods
TL;DR: This article proposes several criteria which isolate specific aspects of the performance of a method, such as its retrieval of inherent structure, its sensitivity to resampling and the stability of its results in the light of new data.
BIRCH: an efficient data clustering method for very large databases
Tian Zhang,Raghu Ramakrishnan,Miron Livny +2 more
- 01 Jun 1996
TL;DR: Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH) as discussed by the authors is a data clustering method that is especially suitable for very large databases.
Computational cluster validation in post-genomic data analysis
TL;DR: In this article, the authors present a review of clustering validation techniques for post-genomic data analysis, with a particular focus on their application to postgenomic analysis of biological data.
993
Evaluation and comparison of gene clustering methods in microarray analysis
TL;DR: The results show that tight clustering and model-based clustering consistently outperform other clustering methods both in simulated and real data while hierarchical clusters and SOM perform among the worst.