Breadth-First Search Crawling Yields High-Quality Pages
Published 1 January 2001
Marc Najork, Janet L. Wiener
Citations182
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This paper examines the average page quality over time of pages downloaded during a web crawl of 328 million unique pages. We use the connectivity-based metric PageRank to measure the quality of a page. We show that traversing the web graph in breadth-first search order is a good crawling strategy, as it tends to discover high-quality pages early on in the crawl.
Keywords
Computer SciencePhysics and Astronomy
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
Symposium on Discrete AlgorithmsAuthoritative sources in a hyperlinked environment
1,836 Citations1998Jon Kleinberg
Computer Networks and ISDN SystemsEfficient crawling through URL ordering
840 Citations1998Junghoo Cho, Héctor García-Molina +1 more
This paper studies in what order a crawler should visit the URLs it has seen, in order to obtain more "important" pages first, and shows that a Crawler with a good ordering scheme can obtain important pages significantly faster than one without.
World Wide WebMercator: A scalable, extensible Web crawler
578 Citations1999Allan Heydon, Marc Najork
This paper describes Mercator, a scalable, extensible Web crawler written entirely in Java, and comments on Mercator's performance, which is found to be comparable to that of other crawlers for which performance numbers have been published.
Computer NetworksOn near-uniform URL sampling
276 Citations2000Monika Henzinger, Allan Heydon +2 more
This paper suggests ways of improving sampling based on random walks of the Web graph to make the samples closer to uniform and suggests a natural test bed based onrandom graphs for testing the effectiveness of the procedures.
Computer Networks and ISDN SystemsThe Connectivity Server: fast access to linkage information on the Web
187 Citations1998Krishna Bharat, Andrei Broder +3 more
A server that provides linkage information for all pages indexed by the AltaVista search engine and can produce the entire neighbourhood of L up to a given distance, and envisage numerous other applications such as ranking, visualization, and classification.
