login

Chinese Word Segmentation Based on Contextual Entropy

Waseda University Repository (Waseda University)Published 1 October 2003
Jin Huang, David Powers
Citations29

TL;DR

This paper presents a new statistical approach to segment Chinese sequences into words based on contextual entropy on both sides of a bigram to capture the dependency with the left and right contexts in which a bigram occurs.

Abstract

Chinese is written without word delimiters so word segmentation is generally considered a key step in processing Chinese texts. This paper presents a new statistical approach to segment Chinese sequences into words based on contextual entropy on both sides of a bigram. It is used to capture the dependency with the left and right contexts in which a bigram occurs. Our approach tries to segment by finding the word boundaries instead of the words. Experimental results show that it is effective for Chinese word segmentation.

Keywords

Computer Science