Forming Word Classes by Statistical Clustering for Statistical Language Modelling
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper presents a clustering algorithm that builds abstract word equivalence classes that finds a local optimum according to a maximum-likelihood criterion and shows good results depending on the size of the training material.
Abstract
In statistical language modelling there is always a problem of sparse data. A way to reduce this problem is to form groups of words in order to get equivalence classes. In this paper we present a clustering algorithm that builds abstract word equivalence classes. The algorithm finds a local optimum according to a maximum-likelihood criterion. Experiments were made on an English 1.1-million word corpus and a German 100,000-word corpus. Compared to a word bigram model, the use of clustered equivalence classes in a bigram class model leads to a significant improvement, as measured by the perplexity. Depending on the size of the training material, the automatically clustered word classes are even better than manually determined categories.KeywordsCluster AlgorithmFrequent WordStatistical ClusterSpeech Recognition SystemWord ClassThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
