Shimmering Lexical Sets
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper shows how coercion phenomena are encoded in the Pattern Dictionary of English Verbs (Hanks and Pustejovsky, 2005), and shows what the CPA Shallow Ontology looks like in its new form, how it is populated, and explores its advantages in terms of empirical validity over more homogeneous, speculative ontologies.
Abstract
0. Abstract For word sense disambiguation and other pedagogical and NLP applications, it has long seemed desirable to group words together according to their essential semantic type 1 — [[Human]], [[Animate]], [[Artefact]], [[Physical Object]], [[Event]], etc.—and to arrange these types into a hierarchy. Vast lexical and conceptual ontologies such as WordNet (Miller and Fellbaum 2007) and the Brandeis Semantic Ontology (Pustejovsky et. al. 2006) have been built on this foundation. However, the expectation that semantic types can serve word sense disambiguation purposes is disappointed by the fact that, as corpus-driven pattern analysis shows, semantic types do not map neatly onto lexical sets (Hanks et al. 2007). Firstly, lexical sets that pick out a particular sense of a verb may cut across semantic types. Secondly, as lexical sets move from verb to verb, some words drop out and others come in. In this paper we examine the implications of these inconvenient observations, and ask how lexical sets of co-occurring words can be organized for purposes of predicting the meaning of words in context. We discuss two steps aimed at dealing with this problem. Firstly, a new type of shallow ontology of nouns is being developed, based on work that was first reported in Pustejovsky, Hanks, and Rumshisky (2004). Secondly, each semantic type is populated by a set of canonical lexical items, which are identified by statistical contextual information. Noncanonical lexical items are classed as exploitations. Exploitations include words that are coerced into honorary membership of a semantic type in particular contexts. In this paper, we show how coercion phenomena are encoded in the Pattern Dictionary of English Verbs (Hanks and Pustejovsky, 2005). We show what the CPA Shallow Ontology looks like in its new form, discuss how it is populated, and explore its advantages in terms of empirical validity over more homogeneous, speculative ontologies.
