login

A stochastic finite-state word-segmentation algorithm for Chinese

Published 1 January 1994Open access
Richard Sproat, William A. Gale, Chilin Shih, Nancy Chang
Citations290
View PDF

TL;DR

A stochastic finite-state model is presented for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names.

Abstract

We present a stochastic finite-state model for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names. We also evaluate the system's performance, taking into account the fact that people often do not agree on a single segmentation.

Keywords

Computer Science