login

A web crawler design for data mining

Journal of Information SciencePublished 1 October 2001
Mike Thelwall
Citations144
SJR quartileQ1
SJR score0.63
SNIP1.31

TL;DR

A distributed design is proposed in order to effectively use idle computing resources and to help information scientists avoid the need to employ dedicated equipment.

Abstract

The content of the web has increasingly become a focus for academic research. Computer programs are needed in order to conduct any large-scale processing of web pages, requiring the use of a web crawler at some stage in order to fetch the pages to be analysed. The processing of the text of web pages in order to extract information can be expensive in terms of processor time. Consequently a distributed design is proposed in order to effectively use idle computing resources and to help information scientists avoid the need to employ dedicated equipment. A system developed using the model is examined and the advantages and limitations of the approach are discussed.

Keywords

Computer Science