Tagset reduction without information loss
Published 1 January 1995Open access
Thorsten Brants
Citations12
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A technique for reducing a tagset used for n-gram part-of-speech disambiguation is introduced and evaluated in an experiment and ensures that all information that is provided by the original tagset can be restored from the reduced one.
Abstract
A technique for reducing a tagset used for n-gram part-of-speech disambiguation is introduced and evaluated in an experiment.
Keywords
Computer Science
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
A stochastic parts program and noun phrase parser for unrestricted text
974 Citations1988Kenneth Church
A program that tags each word in an input sentence with the most likely part of speech has been written and performance is encouraging; a 400-word sample is presented and is judged to be 99.5% correct.
A practical part-of-speech tagger
632 Citations1992Doug Cutting, Julian Kupiec +2 more
An implementation of a part-of-speech tagger based on a hidden Markov model that enables robust and accurate tagging with few resource requirements and accuracy exceeds 96%.
English for the Computer
127 Citations1995Geoffrey Sampson
ArXiv.orgBest-first Model Merging for Hidden Markov Model Induction
115 Citations1994Andreas Stolcke, Stephen M. Omohundro
A new technique for inducing the structure of Hidden Markov Models from data which is based on the general `model merging' strategy, and how the algorithm was incorporated in an operational speech understanding system, where it was combined with neural network acoustic likelihood estimators to improve performance over single-pronunciation word models.
