login

Adding more languages improves unsupervised multilingual part-of-speech tagging

Published 1 January 2009Open access
Benjamin Snyder, Tahira Naseem, Jacob Eisenstein, Regina Barzilay
Citations37
View PDF

TL;DR

A non-parametric Bayesian model is proposed that connects related tagging decisions across languages through the use of multilingual latent variables and shows that performance improves steadily as the number of languages increases.

Abstract

We investigate the problem of unsupervised part-of-speech tagging when raw parallel data is available in a large number of languages. Patterns of ambiguity vary greatly across languages and therefore even unannotated multilingual data can serve as a learning signal. We propose a non-parametric Bayesian model that connects related tagging decisions across languages through the use of multilingual latent variables. Our experiments show that performance improves steadily as the number of languages increases.

Keywords

Computer Science