login

Towards Automatic Web Genre Identification

Hawaii International Conference on System SciencesPublished 7 January 2002
Georg Rehm
Citations43

TL;DR

A database-driven corpus is developed, currently containing 1300000+ documents, which comprises the empirical research basis and the notions of Web genre type which constitutes the framework for a certain Web genre, and compulsory and optional Web genre modules are introduced.

Abstract

We analyse academic Web pages in order to automatically classify them into Web genres. For this purpose, we have developed a database-driven corpus, currently containing 1300000+ documents, which comprises our empirical research basis. We introduce the notions of Web genre type which constitutes the framework for a certain Web genre, and compulsory and optional Web genre modules. These act as building blocks which go together to make up the structure characterised by the Web genre type and operate as modifiers for the default assignment. The analysis of a 200 document sample illustrates our notion of Web genre hierarchy into which Web genre types and modules are embedded. The analysis of four documents of the Web Genre Academic's Personal Homepage demonstrates our approach and our long-term goal of automatically extracting the contents of Web genre modules in order to build up structured XML documents of unstructured HTML documents.

Keywords

Computer Science