login

Lexicalized hidden Markov models for part-of-speech tagging

Published 1 January 2000Open access
Sang-Zoo Lee, Jun’ichi Tsujii, Hae‐Chang Rim
Citations31
View PDF

TL;DR

This paper introduces uniformly lexicalized HMMs for part-of-speech tagging in both English and Korean and uses a simplified back-off smoothing technique to overcome data sparseness.

Abstract

Since most previous works for HMM-based tagging consider only part-of-speech information in contexts, their models cannot utilize lexical information which is crucial for resolving some morphological ambiguity. In this paper we introduce uniformly lexicalized HMMs for part-of-speech tagging in both English and Korean. The lexicalized models use a simplified back-off smoothing technique to overcome data sparseness. In experiments, lexicalized models achieve higher accuracy than non-lexicalized models and the back-off smoothing method mitigates data sparseness better than simple smoothing methods.

Keywords

Computer Science