Building a Large Annotated Corpus of English: The Penn Treebank
Published 30 April 1993
Mitchell P. Marcus
Citations7,528
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
As a result of this grant, the researchers have now published oil CDROM a corpus of over 4 million words of running text annotated with part-of- speech (POS) tags, with over 3 million words of that material assigned skeletal grammatical structure. This material now includes a fully hand-parsed version of the classic Brown corpus. About one half of the papers at the ACL Workshop on Using Large Text Corpora this past summer were based on the materials generated by this grant.
Keywords
Computer Science
ScholarlyCommons (University of Pennsylvania)Part-of-Speech Tagging Guidelines for the Penn Treebank Project (3rd Revision)
436 Citations1990Beatrice Santorini
This manual addresses the linguistic issues that arise in connection with annotating texts by part of speech ("tagging") and discusses parts of speech that are easily confused and gives guidelines on how to tag such cases.
Inside-outside reestimation from partially bracketed corpora
289 Citations1992Fernando Pereira, Yves Schabes
The inside-outside algorithm for inferring the parameters of a stochastic context-free grammar is extended to take advantage of constituent information in a partially parsed corpus to achieve faster convergence and better modelling of hierarchical structure than the original one.
International Conference on Acoustics, Speech, and Signal ProcessingA stochastic parts program and noun phrase parser for unrestricted text
127 Citations2003Kenneth Church
Parsing a natural language using mutual information statistics
127 Citations1990David M. Magerman, Mitchell P. Marcus
The generalized mutual information statistic is derived, the parsing algorithm is described, and results and sample output from the parser are presented.
Acquiring disambiguation rules from text
102 Citations1989Donald Hindle
An effective procedure for automatically acquiring a new set of disambiguation rules for an existing deterministic parser on the basis of tagged text is presented and suggests a path toward more robust and comprehensive syntactic analyzers.
Deducing linguistic structure from the statistics of large corpora
63 Citations1990Eric Brill, David M. Magerman +2 more
Studies in part of speech labelling
24 Citations1991Marie Meteer, Richard Schwartz +1 more
This paper reports experiments in three important areas: handling unknown words, limiting the size of the training set, and returning a set of the most likely tags for each word rather than a single tag.
Discovering the lexical features of a language
22 Citations1991Eric Brill
Evidence is presented for the sufficiency of a strictly syntax-based model for discovering lexical features of a language and whether syntactic cues may not just aid in feature discovery, but may be all that is necessary.
Partial parsing
21 Citations1991Ralph Weischedel, Damaris Ayuso +4 more
This paper reports a handful of experiments designed to test the feasibility of applying well-known partial parsing techniques to the problem of automatic data base update from an open-ended source of messages.
Probabilistic parse scoring based on prosodic phrasing
4 Citations1992Nanette Veilleux, Mari Ostendorf
A decision tree is designed to predict prosodic phrase structure for a given syntactic parse, and the tree is used to compute a parse score, which now is the probability of the recognized break sequence.
