login

Link contexts in classifier-guided topical crawlers

IEEE Transactions on Knowledge and Data EngineeringPublished 1 January 2006
Gautam Pant, Padmini Srinivasan
Citations106
SJR quartileQ1
SJR score2.57
SNIP3.30

TL;DR

This study uses topical crawlers that are guided by a support vector machine to investigate the effects of various definitions of link contexts on the crawling performance and finds that a crawler that exploits words both in the immediate vicinity of a hyperlink as well as the entire parent page performs significantly better than a Crawler that depends on just one of those cues.

Abstract

Context of a hyperlink or link context is defined as the terms that appear in the text around a hyperlink within a Web page. Link contexts have been applied to a variety of Web information retrieval and categorization tasks. Topical or focused Web crawlers have a special reliance on link contexts. These crawlers automatically navigate the hyperlinked structure of the Web while using link contexts to predict the benefit of following the corresponding hyperlinks with respect to some initiating topic or theme. Using topical crawlers that are guided by a support vector machine, we investigate the effects of various definitions of link contexts on the crawling performance. We find that a crawler that exploits words both in the immediate vicinity of a hyperlink as well as the entire parent page performs significantly better than a crawler that depends on just one of those cues. Also, we find that a crawler that uses the tag tree hierarchy within Web pages provides effective coverage. We analyze our results along various dimensions such as link context quality, topic difficulty, length of crawl, training data, and topic domain. The study was done using multiple crawls over 100 topics covering millions of pages allowing us to derive statistically strong results.

Keywords

Computer Science