Pagerank based clustering of hypertext document collections
Published 20 July 2008
Konstantin Avrachenkov, Vladimir Dobrynin, Danil Nemirovsky, Son Pham, Elena Smirnova
Citations50
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A novel PageRank based clustering (PRC) algorithm which uses the hypertext structure is proposed which produces graph partitioning with high modularity and coverage.
Abstract
International audience
Keywords
Computer SciencePhysics and Astronomy
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
Proceedings of the National Academy of SciencesModularity and community structure in networks
12,281 Citations2006M. E. J. Newman
It is shown that the modularity of a network can be expressed in terms of the eigenvectors of a characteristic matrix for the network, which is called modularity matrix, and that this expression leads to a spectral algorithm for community detection that returns results of demonstrably higher quality than competing methods in shorter running times.
Graph Clustering by Flow Simulation
1,608 Citations2000S.M. van Dongen
Topic-sensitive PageRank
1,524 Citations2002Taher H. Haveliwala
A set of PageRank vectors are proposed, biased using a set of representative topics, to capture more accurately the notion of importance with respect to a particular topic, and are shown to generate more accurate rankings than with a single, generic PageRank vector.
Inferring Web communities from link topology
772 Citations1998David Gibson, Jon Kleinberg +1 more
This investigation shows that although the process by which users of the Web create pages and links is very difficult to understand at a “local” level, it results in a much greater degree of orderly high-level structure than has typically been assumed.
Lecture notes in computer scienceExperiments on Graph Clustering Algorithms
303 Citations2003Ulrik Brandes, Marco Gaertler +1 more
An experimental evaluation of graph clustering approaches is conducted and by combining proven techniques from graph partitioning and geometric clustering, a new approach is introduced that compares favorably.
Clustering by weighted cuts in directed graphs
129 Citations2007Marina Meilă, William Pentney
It is shown that this problem can be relaxed to a Rayleigh quotient problem for a symmetric matrix obtained from the original affinities and therefore a large body of the results and algorithms developed for spectral clustering of symmetric data immediately extends to asymmetric cuts.
Internet MathematicsLocal Partitioning for Directed Graphs Using PageRank
61 Citations2008Reid Andersen, Fan Chung +1 more
It is proved that by computing a personalized PageRank vector in a directed graph, starting from a single seed vertex within a set S that has conductance at most α, and by performing a sweep over that vector, one can obtain a set of vertices S′ with conductance.
Lecture notes in computer scienceWeb Communities Identification from Random Walks
30 Citations2006Jiayuan Huang, Tingshao Zhu +1 more
To efficiently extract communities from the stationary distribution defined by a random walk, a computationally efficient form of directed spectral clustering is exploited, which is shown to effectively identify latent Web communities based on link topology only.
Lecture notes in computer scienceContextual Document Clustering
30 Citations2004Vladimir Dobrynin, David Patterson +1 more
A novel algorithm based on distributional clustering where subject related words, which have a narrow context, are identified to form meta-tags for that subject to form thematic clusters of documents is presented.
