login

Automatic construction of polarity-tagged corpus from HTML documents

Published 1 January 2006Open access
Nobuhiro Kaji, Masaru Kitsuregawa
Citations92
View PDF

TL;DR

A novel method of building polarity-tagged corpus from HTML documents that can be applied to arbitrary HTML documents and utilize certain layout structures and linguistic pattern to automatically extract sentences that express opinion.

Abstract

This paper proposes a novel method of building polarity-tagged corpus from HTML documents. The characteristics of this method is that it is fully automatic and can be applied to arbitrary HTML documents. The idea behind our method is to utilize certain layout structures and linguistic pattern. By using them, we can automatically extract such sentences that express opinion. In our experiment, the method could construct a corpus consisting of 126,610 sentences.

Keywords

Computer Science