login

Discrimination decisions for 100,000-dimensional spaces

Annals of Operations ResearchPublished 1 June 1995
William A. Gale, Kenneth Church, David Yarowsky
Citations13
SJR quartileQ1
SJR score1.09
SNIP1.62

TL;DR

A method designed for the sense discrimination problems mentioned is introduced and areas for research based on observed shortcomings of the method are discussed, including the need for a robust version of this method.

Abstract

Discrimination decisions arise in many natural language processing tasks. Three classical tasks are discriminating texts by their authors (author identification), discriminating documents by their relevance to some query (information retrieval), and discriminating multi-meaning words by their meanings (sense discrimination). Many other discrimination tasks arise regularly, such as determining whether a particular proper noun represents a person or a place, or whether a given work from some teletype text would be capitalized if both cases had been used. We [9] introduced a method designed for the sense discrimination problems mentioned. We also discuss areas for research based on observed shortcomings of the method. In particular, an example in the author identification task shows the need for a robust version of the method. Also, the method makes an assumption of independence which is demonstrably false, yet there has been no careful study of the results of this assumption.

Keywords

Computer Science