Experiments in Unsupervised Entropy-Based Corpus Segmentation
Published 1 January 1999
André Kempe
Citations20
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
The paper presents an entropy-based approach to segment a corpus into words, when no additional information about the corpus or the language, and no other resources such as a lexicon or grammar are available. To segment the corpus, the algorithm searches for separators, without knowing a priori by which symbols they are constituted. Good results can be obtained with corpora containing "clearly perceptible" separators such as blank or new-line.
Keywords
Computer Science
Bell System Technical JournalA Mathematical Theory of Communication
9,723 Citations1948Claude E. Shannon
From Phoneme to Morpheme
303 Citations1970Zellig S. Harris
British Journal of PsychologyThe discovery of segments in natural language
57 Citations1977J. Gerard Wolff
The performance of a computer model of linguistic segmentation is described and evaluated when it is used with natural language and seems to have some sensitivity to morphs but it performs poorly with structures larger than words.
Finding structure via compression
17 Citations1998Jason Hutchens, M. Alder
The results show that language models which bootstrap themselves with structure found in this way undergo a reduction in perplexity, and it is concluded that these techniques may be useful in the design of generic grammatical inference systems.
