Three Generative, Lexicalised Models for Statistical Parsing
arXiv (Cornell University)Published 17 June 1997Open access
Michael J. Collins
Citations50
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In this paper we first propose a new statistical parsing model, which is a generative model of lexicalised context-free grammar. We then extend the model to include a probabilistic treatment of both subcategorisation and wh-movement. Results on Wall Street Journal text show that the parser performs at 88.1/87.5% constituent precision/recall, an average improvement of 2.3% over (Collins 96).
Keywords
Computer Science
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
The Philosophical ReviewGeneralized Phrase Structure Grammar.
1,843 Citations1989Scott Soames, Gerald Gazdar +3 more
"Generalized Phrase Structure Grammar" provides the definitive exposition of the theory of grammar originally proposed by Gerald Gazdar and developed during half a dozen years' work with his colleagues Ewan Klein, Geoffrey Pullum, and Ivan Sag.
Three new probabilistic models for dependency parsing
630 Citations1996Jason Eisner
Preliminary empirical results from evaluating the three models' parsing performance on annotated Wall Street Journal training text (derived from the Penn Treebank) suggest the generative model performs significantly better than the others, and does about equally well at assigning part-of-speech tags.
A new statistical parser based on bigram lexical dependencies
619 Citations1996Michael Collins
A new statistical parser which is based on probabilities of dependencies between head-words in the parse tree, which trains on 40,000 sentences in under 15 minutes and can be improved to over 200 sentences a minute with negligible loss in accuracy.
IEEE Transactions on ComputersApplying Probability Measures to Abstract Languages
232 Citations1973Taylor L. Booth, Richard A. Thompson
The problem of assigning a probability to each word of a language is considered and two methods are discussed.
A fully statistical approach to natural language interfaces
143 Citations1996S.L. Miller, David Stallard +2 more
This work presents a natural language interface system which is based entirely on trained statistical models, resulting in an end-to-end system that maps input utterances into meaning representation frames.
Decision tree parsing using a hidden derivation model
74 Citations1994F. Jelinek, John Lafferty +4 more
The grammarian then evaluates the performance of the grammar, and upon analysis of the errors made by the grammar-based parser, carefully refines the rules, repeating this process, typically over a period of several years.
arXiv (Cornell University)Statistical Decision-Tree Models for Parsing
46 Citations1995David M. Magerman
Parsing with Context-Free Grammars and Word Statistics
30 Citations1995Eugene Charniak
A language model in which the probability of a sentence is the sum of the individual parse probabilities, and these are calculated using a probabilistic context-free grammar (PCFG) plus statistics on individual words and how they fit into parses is presented.
