Unsupervised segmentation of Chinese text by use of branching entropy
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
An unsupervised segmentation method based on an assumption that the increasing point of entropy of successive characters is the location of a word boundary is proposed, finding that the precision was stable at around 90% independently of the learning data size.
Abstract
We propose an unsupervised segmentation method based on an assumption about language data: that the increasing point of entropy of successive characters is the location of a word boundary. A large-scale experiment was conducted by using 200 MB of unsegmented training data and 1 MB of test data, and precision of 90% was attained with recall being around 80%. Moreover, we found that the precision was stable at around 90% independently of the learning data size.
