High-precision identification of discourse new and unique noun phrases
Published 1 January 2003Open access
Olga Uryupina
Citations37
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This study tries to learn automatically two classifications, ±discourse_new and ±unique, relevant for coreference resolution systems, and expects these classifiers to provide a good prefiltering for core Conference Resolution systems, improving both their speed and performance.
Abstract
Coreference resolution systems usually attempt to find a suitable antecedent for (almost) every noun phrase. Recent studies, however, show that many definite NPs are not anaphoric. The same claim, obviously, holds for the indefinites as well.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
Elsevier eBooksFast Effective Rule Induction
3,767 Citations1995William W. Cohen
This paper evaluates the recently-proposed rule learning algorithm IREP on a large and diverse collection of benchmark problems, and proposes a number of modifications resulting in an algorithm RIPPERk that is very competitive with C4.5 and C 4.5rules with respect to error rates, but much more efficient on large samples.
A maximum-entropy-inspired parser
1,496 Citations2000Eugene Charniak
A new parser for parsing down to Penn tree-bank style parse trees that achieves 90.1% average precision/recall for sentences of length 40 and less and 89.5% when trained and tested on the previously established sections of the Wall Street Journal treebank is presented.
Computational LinguisticsAn Empirically Based System for Processing Definite Descriptions
198 Citations2000Renata Vieira, Massimo Poesio
The annotated corpus was used to extensively evaluate the proposed techniques for matching definite descriptions with their antecedents, discourse segmentation, recognizing discourse-new descriptions, and suggesting anchors for bridging descriptions.
Identifying anaphoric and non-anaphoric noun phrases to improve coreference resolution
158 Citations2002Vincent Ng, Claire Cardie
A supervised learning approach to identification of anaphoric and non-anaphoric noun phrases is presented and it is shown how such information can be incorporated into a coreference resolution system that outperforms the best M UC-6 and MUC-7 coreferenceresolution systems on the corresponding MUC coreference data sets.
Using the web to overcome data sparseness
97 Citations2002Frank Keller, Maria Lapata +1 more
It is shown that the web can be employed to obtain frequencies for bigrams that are unseen in a given corpus by demonstrating that web frequencies and correlate with frequencies obtained from a carefully edited, balanced corpus.
Corpus-based identification of non-anaphoric noun phrases
66 Citations1999David L. Bean, Ellen Riloff
A corpus-based algorithm is developed for automatically identifying definite noun phrases that are non-anaphoric, which has the potential to improve the efficiency and accuracy of coreference resolution systems.
