Language-Independent Parsing with Empty Elements
Published 19 June 2011
Shu Cai, David Chiang, Yoav Goldberg
Citations49
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A simple, language-independent method for integrating recovery of empty elements into syntactic parsing that outperforms the best published method the authors are aware of on English and a recently published method on Chinese.
Abstract
We present a simple, language-independent method for integrating recovery of empty ele-ments into syntactic parsing. This method out-performs the best published method we are aware of on English and a recently published method on Chinese. 1
Keywords
Computer Science
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Learning accurate, compact, and interpretable tree annotation
812 Citations2006Slav Petrov, Leon Barrett +2 more
An automatic approach to tree annotation in which basic nonterminal symbols are alternately split and merged to maximize the likelihood of a training treebank is presented.
Natural Language EngineeringThe Penn Chinese TreeBank: Phrase structure annotation of a large corpus
694 Citations2005Nianwen Xue, Fei Xia +2 more
Several Chinese linguistic issues and their implications for treebanking efforts are discussed and how to address these issues when developing annotation guidelines are addressed, and engineering strategies to improve speed while ensuring annotation quality are described.
Procedure for quantitatively comparing the syntactic coverage of English grammars
436 Citations1991Judith L. Klavans, Mark Liberman +12 more
Better Arabic Parsing: Baselines, Evaluations, and Analysis
176 Citations2010Spence Green, Christopher D. Manning
This paper identifies sources of syntactic ambiguity understudied in the existing parsing literature, and develops a human interpretable grammar that is competitive with a latent variable PCFG.
A simple pattern-matching algorithm for recovering empty nodes and their antecedents
111 Citations2001Mark Johnson
An evaluation procedure for empty node recovery procedures is proposed which is independent of most of the details of phrase structure, which makes it possible to compare the performance of empty nodes recovery on parser output with the empty node annotations in a gold-standard corpus.
Fully parsing the Penn Treebank
79 Citations2006Ryan Gabbard, Mitchell P. Marcus +1 more
A two stage parser that recovers Penn Treebank style syntactic analyses of new sentences including skeletal syntactic structure, and, for the first time, both function tags and empty categories is presented.
A Single Generative Model for Joint Morphological Segmentation and Syntactic Parsing
61 Citations2008Yoav Goldberg, Reut Tsarfaty
A single joint model for performing both morphological segmentation and syntactic disambiguation which bypasses the associated circularity is proposed.
Using linguistic principles to recover empty categories
51 Citations2004Richard Campbell
An algorithm for detecting empty nodes in the Penn Treebank, finding their antecedents, and assigning them function tags, without access to lexical information such as valency is described.
Infoscience (Ecole Polytechnique Fédérale de Lausanne)Lattice Parsing for Speech Recognition
49 Citations1999Martin Rajman, Rosario Aragüés +2 more
The main goal is to present and demonstrate by actual experiments that sequential coupling may be efficiently achieved by word-lattice syntactic analyzers, efficiently parsing the huge num ber of hypothesis (i.e. possible sentences) contained in the lattice produced by the speech recognizer.
Empirical Methods in Natural Language ProcessingEffects of Empty Categories on Machine Translation
46 Citations2010Tagyoung Chung, Daniel Gildea
Different sperm characteristics respond differently at low temperatures and the freezing of buffalo spermatozoa at a higher rate ensures higher post-thaw semen quality.
Trace prediction and recovery with unlexicalized PCFGs and slash features
41 Citations2006Helmut Schmid
A parser which generates parse trees with empty elements in which traces and fillers are co-indexed and which outperformed other unlexicalized PCFG parsers in terms of labeled bracketing f-score.
Antecedent recovery
40 Citations2003Péter Dienes, Amit Dubey
This paper develops both a two- step approach which combines a trace tagger with a state-of-the-art lexicalized parser and a one-step approach which finds nonlocal dependencies while parsing.
Chasing the ghost: recovering empty categories in the Chinese Treebank
39 Citations2010Yaqin Yang, Nianwen Xue
A unified framework in recovering empty categories in the Chinese Tree-bank is described and the results show that given skeletal gold standard parses, the empty categories can be detected with very high accuracy.
Joint Hebrew Segmentation and Parsing using a PCFGLA Lattice Parser
26 Citations2011Yoav Goldberg, Michael Elhadad
This work constructs a parser which parses and segments unsegmented Hebrew text with an F-score of almost 80, an error reduction of over 20% over the best previous result for this task, indicating that lattice parsing with the Berkeley parser is an effective methodology for parsing over uncertain inputs.
Best-first word-lattice parsing: techniques for integrated syntactic language modeling
16 Citations2005Mark Johnson, Keith Hall
This thesis presents a best-first word-lattice chart parsing algorithm which combines the search for good parses with theSearch for good strings in the word- lattice to provide an efficient syntactic language model.
