login

Extracting Key Terms from Chinese and Japanese texts

Published 1 January 1998
Pascale Fung
Citations26

TL;DR

This paper presents two simple yet powerful systems for Chinese and Japanese key term extraction|CXtract and JBrat, which are based on morphosyn-tactic information of the Japanese character sets for terms.

Abstract

Key term extraction is very useful for information retrieval. Most term extraction methods use one of two approaches, namely lexical and grammatical. We argue that due to the differences in linguistic and character set characteristics of Chinese and Japanese, a lexical approach is more suitable for Chinese whereas a grammatical approach is more suitable for Japanese. In this paper, we present two simple yet powerful systems for Chinese and Japanese key term extraction---CXtract and JBrat. CXtract uses predominantly statistical lexical information to find term boundaries in large text. JBrat is based on morphosyntactic information of the Japanese character sets for terms. Evaluation results show that CXtract has a 80.24% average precision in term extraction, and JBrat has a 88.07% average precision. 1 Introduction Linguists have argued that the smallest semantic unit is often not a single word, as defined by a string of letters delimited by spaces, but a phrase (or a term) (Pinchuck 19...

Keywords

Computer Science