login

Chinese Named Entity Recognition combining a statistical model with human knowledge

Published 1 January 2003Open access
Youzheng Wu, Jun Zhao, Bo Xu
Citations52
View PDF

TL;DR

This paper presents a hybrid algorithm which can combine a class-based statistical model with various types of human knowledge very well and employs a back-off model and a Chinese thesaurus, to smooth the parameters in the model.

Abstract

Named Entity Recognition is one of the key techniques in the fields of natural language processing, information retrieval, question answering and so on. Unfortunately, Chinese Named Entity Recognition (NER) is more difficult for the lack of capitalization information and the uncertainty in word segmentation. In this paper, we present a hybrid algorithm which can combine a class-based statistical model with various types of human knowledge very well. In order to avoid data sparseness problem, we employ a back-off model and [Abstract contained text which could not be captured.], a Chinese thesaurus, to smooth the parameters in the model. The F-measure of person names, location names, and organization names on the newswire test data for the 1999 IEER evaluation in Mandarin is 86.84%, 84.40% and 76.22% respectively.

Keywords

Computer Science