login

Exploiting Structural Information for Text Classification on the WWW

Lecture notes in computer sciencePublished 1 January 1999
Johannes Fürnkranz
Citations123
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

Experimental evidence is presented that confirms the working hypothesis that it is often easier to classify a hypertext page using information provided on pages that point to it instead of using information that is provided on the page itself.

Abstract

In this paper, we report on a set of experiments that explore the utility of making use of the structural information of WWW documents. Our working hypothesis is that it is often easier to classify a hypertext page using information provided on pages that point to it instead of using information that is provided on the page itself. We present experimental evidence that confirms this hypothesis on a set of Web-pages that relate to Computer Science Departments.

Keywords

Computer Science