login

Unsupervised segmentation of Chinese text by use of branching entropy

Published 1 January 2006Open access
Zhihui Jin, Kumiko Tanaka‐Ishii
Citations66
View PDF

TL;DR

An unsupervised segmentation method based on an assumption that the increasing point of entropy of successive characters is the location of a word boundary is proposed, finding that the precision was stable at around 90% independently of the learning data size.

Abstract

We propose an unsupervised segmentation method based on an assumption about language data: that the increasing point of entropy of successive characters is the location of a word boundary. A large-scale experiment was conducted by using 200 MB of unsegmented training data and 1 MB of test data, and precision of 90% was attained with recall being around 80%. Moreover, we found that the precision was stable at around 90% independently of the learning data size.

Keywords

Computer Science