login

Proper Name Extraction from Non-Journalistic Texts

Published 1 January 2001
Thierry Poibeau, Leila Kosseim
Citations84

TL;DR

The influence of the corpus on the automatic identification of proper names in texts is discussed and an approach to adapt a proper name extraction system developed for newspapers to the analysis of e-mail is described.

Abstract

In this paper, the syntactic properties of parenthetical reporting clauses in Dutch are investigated by means of a small corpus. It is shown that an analysis in which the quote is looked upon as the direct object of the reporting verb is inadequate. Therefore, an alternative analysis is proposed, viz. one in which the quote and the reporting clause are taken to be adjoined. Such an analysis, however, is not unproblematic. Here we discuss the two main problems with this analysis.

Keywords

Computer Science