login

Using Grammatical Features for Automatic Register Identification in an Unrestricted Corpus of Documents from the Open Web

Journal of Research Design and Statistics in Linguistics and Communication SciencePublished 16 February 2016
Douglas Biber, Jesse Egbert
Citations50

TL;DR

The findings demonstrate the possibility of automatically predicting register/genre on the unrestricted open web, and it is anticipated that future extensions will allow this task to be accomplished with considerably higher degrees of accuracy.

Abstract

Most previous attempts at automatic genre identification have been based on corpus samples that are relatively small and artificially restricted. In this study we set out to automatically predict register/genre categories in a large, representative sample of documents from the open web using a linguistic approach focused on lexico-grammatical characteristics that have functional associations. Our findings demonstrate the possibility of automatically predicting register/genre on the unrestricted open web, and we anticipate that future extensions will allow this task to be accomplished with considerably higher degrees of accuracy.

Keywords

Computer Science