Chinese word sense disambiguation by combining pseudo training data
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A kind of substitute for Chinese sense tagged data, which is called pseudo training data, is suggested, based on a linguistic phenomenon in Chinese, which some multicharacter-words inherit only one sense and some syntactical features from ambiguous one-character-words.
Abstract
In supervised methods of word sense disambiguation, sense tagged samples for training classifiers are needed. Since sense tagging for ambiguous words is expensive and labor intensive, it is worth looking for some reasonable substitutes. We suggest a kind of substitute for Chinese sense tagged data, which is called pseudo training data. The suggestion is based on a linguistic phenomenon in Chinese, which some multicharacter-words inherit only one sense and some syntactical features from ambiguous one-character-words. Data derived from unambiguous multicharacter-words are employed as pseudo training data. Pseudo training data have an advantage of being able to be collected automatically. Our experiments show that classifiers trained by not too much pseudo training data outperform classifiers trained by small quantities of sense tagged samples for Chinese ambiguous word senses. Further experiments show combination of a small set of tagged data and a large quantities pseudo training data is a more promising way to word sense disambiguation.
