Testing Graph Clusterability: Algorithms and Lower Bounds
Ashish Chiplunkar,Michael Kapralov,Sanjeev Khanna,Aida Mousavifar,Yuval Peres +4 more
- 01 Oct 2018
- pp 497-508
TL;DR: In this article, Czumaj, Peng, and Sohler gave a sublinear time algorithm for testing k-clusterability in time O(n^1/2 poly(k)).
read more
Abstract: We consider the problem of testing graph cluster structure: given access to a graph G = (V, E), can we quickly determine whether the graph can be partitioned into a few clusters with good inner conductance, or is far from any such graph? This is a generalization of the well-studied problem of testing graph expansion, where one wants to distinguish between the graph having good expansion (i.e. being a good single cluster) and the graph having a sparse cut (i.e. being a union of at least two clusters). A recent work of Czumaj, Peng, and Sohler (STOC'15) gave an ingenious sublinear time algorithm for testing k-clusterability in time O(n^1/2 poly(k)). Their algorithm implicitly embeds a random sample of vertices of the graph into Euclidean space, and then clusters the samples based on estimates of Euclidean distances between the points. This yields a very efficient testing algorithm, but only works if the cluster structure is very strong: it is necessary to assume that the gap between conductances of accepted and rejected graphs is at least logarithmic in the size of the graph G. In this paper we show how one can leverage more refined geometric information, namely angles as opposed to distances, to obtain a sublinear time tester that works even when the gap is a sufficiently large constant. Our tester is based on the singular value decomposition of a natural matrix derived from random walk transition probabilities from a small sample of seed nodes. We complement our algorithm with a matching lower bound on the query complexity of testing clusterability. Our lower bound is based on a novel property testing problem, which we analyze using Fourier analytic tools. As a byproduct of our techniques, we also achieve new lower bounds for the problem of approximating MAX-CUT value in sublinear time.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Testing Graph Clusterability: Algorithms and Lower Bounds
Ashish Chiplunkar,Michael Kapralov,Sanjeev Khanna,Aida Mousavifar,Yuval Peres +4 more
- 01 Oct 2018
TL;DR: In this article, Czumaj, Peng, and Sohler gave a sublinear time algorithm for testing k-clusterability in time O(n^1/2 poly(k)).
19
•Posted Content
Walking Randomly, Massively, and Efficiently
TL;DR: In this paper, a set of techniques that allow for efficiently generating many independent random walks in the MPC model with space per machine strongly sublinear in the number of vertices is introduced.
12
•Proceedings Article
Robust clustering oracle and local reconstructor of cluster structure of graphs
Pan Peng
- 05 Jan 2020
TL;DR: This work formalizes the notion of robust clustering oracle for a noisy clusterable graph, and gives an algorithm that builds such an oracle in sublinear time, which can be further used to support typical queries regarding the cluster structure of the graph in sub linear time.
12
A Statistical Test of Heterogeneous Subgraph Densities to Assess Clusterability
Pierre Miasnikof,Liudmila Ostroumova Prokhorenkova,Alexander Y. Shestopaloff,Andrei Mikhailovich Raigorodskii +3 more
- 27 May 2019
TL;DR: A novel statistical test, the \(\delta \)-test, which is based on comparisons of local and global densities is introduced, to assess whether a given graph meets the necessary conditions to be meaningfully summarized by clusters of vertices.
•Posted Content
Graph Clustering Via QUBO and Digital Annealing.
TL;DR: This article empirically examines the computational cost of solving a known hard problem, graph clustering, using novel purpose-built computer hardware using a quadratic unconstrained binary optimization problem and employs a novel computer architecture to obtain a numerical solution.
3
References
A Probabilistic Proof of an Asymptotic Formula for the Number of Labelled Regular Graphs
TL;DR: The method determines the asymptotic distribution of the number of short cycles in graphs with a given degree sequence, and gives analogous formulae for hypergraphs.
1.4K
On clusterings-good, bad and spectral
Ravi Kannan,Santosh Vempala,A. Veta +2 more
- 12 Nov 2000
TL;DR: Two results regarding the quality of the clustering found by a popular spectral algorithm are presented, one proffers worst case guarantees whilst the other shows that if there exists a "good" clustering then the spectral algorithm will find one close to it.
1K
On clusterings: Good, bad and spectral
TL;DR: A natural bicriteria measure for assessing the quality of a clustering that avoids the drawbacks of existing measures is motivated and a simple recursive heuristic is shown to have poly-logarithmic worst-case guarantees under the new measure.
906
The isoperimetric number of random regular graphs
TL;DR: It is shown that for every e > 0 there is a natural number r such that in almost every r-regular graph of order n, every set of u ≤ n/2 vertices is joined by at least (r/2 - e)u edges to the rest of the graph.
264
•Journal Article
Testing Expansion in Bounded Degree Graphs
Satyen Kale,C. Seshadhri +1 more
TL;DR: A property tester is given that given a graph with degree bound d, an expansion bound �, and a parameter " > 0, accepts the graph with high probability if its expansion is more than�, and rejects it withhigh probability if it is "-far from agraph with expansion � 0 withdegree bound d".
172
Related Papers (5)
Artur Czumaj,Pan Peng,Christian Sohler +2 more
- 14 Jun 2015
Lorenzo Orecchia,Zeyuan Allen-Zhu +1 more
- 05 Jan 2014