Link Analysis in Web Information Retrieval.
Published 1 January 2000
Monika Henzinger
Citations111
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This survey describes two successful link analysis algorithms and the state-of-the art of the field.
Abstract
The analysis of the hyperlink structure of the web has led to significant improvements in web information retrieval. This survey describes two successful link analysis algorithms and the state-of-the art of the field.
Keywords
Computer SciencePhysics and Astronomy
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
Modern Information Retrieval
11,544 Citations1999Ricardo Baeza‐Yates, Berthier Ribeiro‐Neto
Journal of the American Society for Information ScienceCo‐citation in the scientific literature: A new measure of the relationship between two documents
5,183 Citations1973Henry Small
A new form of document coupling called co-citation is defined as the frequency with which two documents are cited together, and clusters of co- cited papers provide a new way to study the specialty structure of science.
PsychometrikaA New Status Index Derived from Sociometric Analysis
3,863 Citations1953Leo Katz
A new method of computation which takes into account who chooses as well as how many choose is presented, which introduces the concept of attenuation in influence transmitted through intermediaries.
American DocumentationBibliographic coupling between scientific papers
2,765 Citations1963M. M. Kessler
The population of papers under study was ordered into groups that satisfy the stated criterion of interrelation and an examination of the papers that constitute the groups shows a high degree of logical correlation.
ScienceCitation Analysis as a Tool in Journal Evaluation
2,760 Citations1972Eugene Garfield
In 1971, the Institute for Scientfic Information decided to undertake a systematic analysis of journal citation patterns across the whole of science and technology.
Symposium on Discrete AlgorithmsAuthoritative sources in a hyperlinked environment
1,836 Citations1998Jon Kleinberg
Lecture notes in computer scienceThe Web as a Graph: Measurements, Models, and Methods
1,009 Citations1999Jon Kleinberg, Ravi Kumar +3 more
This paper describes two algorithms that operate on the Web graph, addressing problems from Web search and automatic community discovery, and proposes a new family of random graph models that point to a rich new sub-field of the study of random graphs, and raises questions about the analysis of graph algorithms on the Internet.
Computer Networks and ISDN SystemsEfficient crawling through URL ordering
840 Citations1998Junghoo Cho, Héctor García-Molina +1 more
This paper studies in what order a crawler should visit the URLs it has seen, in order to obtain more "important" pages first, and shows that a Crawler with a good ordering scheme can obtain important pages significantly faster than one without.
Enhanced hypertext categorization using hyperlinks
775 Citations1998Soumen Chakrabarti, Byron Dom +1 more
This work has developed a text classifier that misclassified only 13% of the documents in the well-known Reuters benchmark; this was comparable to the best results ever obtained and its technique also adapts gracefully to the fraction of neighboring documents having known topics.
Computer Networks and ISDN SystemsAutomatic resource compilation by analyzing hyperlink structure and associated text
700 Citations1998Soumen Chakrabarti, Byron Dom +4 more
An evaluation of ARC suggests that the resources found by ARC frequently fare almost as well as, and sometimes better than, lists of resources that are manually compiled or classified into a topic.
Computer NetworksLink prediction and path analysis using Markov chains
508 Citations2000Ramesh R. Sarukkai
The generality and power of Markov chains is a first step towards the application of powerful probabilistic models to Web path analysis and link prediction.
Computer NetworksThe stochastic approach for link-structure analysis (SALSA) and the TKC effect
480 Citations2000Ronny Lempel, Shlomo Moran
SALSA, a new stochastic approach for link structure analysis, which examines random walks on graphs derived from the link structure, is presented and it is proved that SALSA is equivalent to a weighted in-degree analysis of the link-structure of World Wide Web subgraphs, making it computationally more efficient than the mutual reinforcement approach.
Efficient Computation of PageRank
294 Citations1999Taher H. Haveliwala
It is shown that PageRank can be computed for very large subgraphs of the web (up to hundreds of millions of nodes) on machines with limited main memory.
Computer NetworksOn near-uniform URL sampling
276 Citations2000Monika Henzinger, Allan Heydon +2 more
This paper suggests ways of improving sampling based on random walks of the Web graph to make the samples closer to uniform and suggests a natural test bed based onrandom graphs for testing the effectiveness of the procedures.
Does “authority” mean quality? predicting expert quality ratings of Web documents
227 Citations2000Brian Amento, Loren Terveen +1 more
An experimental evaluation of link analysis algorithms for their potential to identify high quality items using a dataset of web documents rated for quality by human topic experts found link-based metrics did a good job of picking out high-quality items.
Columbia Academic Commons (Columbia University)Computing Geographical Scopes of Web Resources
225 Citations2000Junyan Ding, Luis Gravano +1 more
Techniques for automatically computing the geographical scope of web resources, based on the textual content of the resources, as well as on the geographical distribution of hyperlinks to them are introduced.
Computer Networks and ISDN SystemsWebQuery: searching and visualizing the Web through connectivity
174 Citations1997S. Jeromy Carrière, Rick Kazman
This work examines links among the nodes returned in a keyword-based query, finding “interesting” sites that are highly connected to those sites returned by the original query by finding ‘hot spots’ on the Web that contain information germane to a user's query.
Exploiting Geographical Location Information of Web Pages
157 Citations1999O. Buyukokkten, J. Cho +3 more
This paper makes the case for identifying and exploiting the geographical location information of web sites so that web search engines can rank resources in a geographically sensitive fashion, in addition to using more traditional information-retrieval strategies.
Computer NetworksWhat is this page known for? Computing Web page reputations
125 Citations2000Davood Rafiei, Alberto O. Mendelzon
A search process where the input is the URL of a page, and the output is a ranked set of topics on which the page has a reputation, and several algorithmic formulations of the notion of reputation are proposed.
Clustering hypertext with applications to web searching
94 Citations2000Dharmendra S. Modha, W. Scott Spangler
A method and structure of searching a database containing hypertext documents comprising searching the database using a query to produce a set ofhypertext documents and geometrically clustering the set ofHypertext documents into various clusters using a toric k-means similarity measure.
