An Empirical Verification of Coverage and Correctness for a General-Purpose Sentence Generator
Published 1 July 2002
Irene Langkilde-Geary
Citations150
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This paper describes a general-purpose sentence generation system that can achieve both broad scale coverage and high quality while aiming to be suitable for a variety of generation tasks. We measure the coverage and correctness empirically using a section of the Penn Treebank corpus as a test set. We also describe novel features that help make the generator flexible and easier to use for a variety of tasks. To our knowledge, this is the first empirical measurement of coverage reported in the literature, and the highest reported measurements of correctness.
Keywords
Computer Science
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Statistical language modeling using the CMU-cambridge toolkit
564 Citations1997Philip Clarkson, Roni Rosenfeld
The conventional language modeling technology, as implemented in the toolkit, is outlined, and the extra e(cid:14)ciency and functionality that the new toolkit provides as compared to previous software for this task is described.
Generation that exploits corpus-based statistical knowledge
355 Citations1998Irene Langkilde, Kevin Knight
Novel aspects of a new natural language generator called Nitrogen are described, which has a highly flexible input representation that allows a spectrum of input from syntactic to semantic depth, and shifts the burden of many linguistic decisions to the statistical post-processor.
Forest-based statistical sentence generation
147 Citations2000Irene Langkilde
A new approach to statistical sentence generation is presented in which alternative phrases are represented as packed sets of trees, or forests, and then ranked statistically to choose the best one, and an efficient ranking algorithm is described.
Evaluation metrics for generation
127 Citations2000Srinivas Bangalore, Owen Rambow +1 more
It is confirmed that intrinsic metrics cannot replace human evaluation, but some correlate significantly with human judgments of quality and understandability and can be used for evaluation during development.
Two-level, many-paths generation
116 Citations1995Kevin Knight, Vasileios Hatzivassiloglou
A hybrid generator is built, in which gaps in symbolic knowledge are filled by statistical methods, to attack problems of large-scale natural language generation and to simplify current generators and enhance their portability.
Columbia Academic Commons (Columbia University)Revision-Based Generation of Natural Language Summaries Providing Historical Background: Corpus-Based Analysis, Design, Implementation and Evaluation
80 Citations1994Jacques Robin
This thesis presents a new generation model in which a first pass builds a draft containing only the essential new facts to report and a second pass incrementally revises this draft to opportunistically add as many background facts as can fit within the space limit.
FUF: the Universal Unifier User Manual Version 2.0
46 Citations1989Michael Elhadad
This document is the user manual for FUF version 5.2, a natural language generator program that uses the technique of unification grammars, and includes novel techniques in the unification allowing the specification of types and the expression of complete information.
Scholarworks (University of Massachusetts Amherst)The “GENERATION GAP”: the problem of expressibility in text planning
44 Citations1990Marie Meteer
This thesis identifies and provides a solution for a particular problem in natural language generation: the problem of ensuring the expressibility of a text plan by designing a level of representation, the Text Structure, which is used by the text planner in composing the utterance.
arXiv (Cornell University)Filling Knowledge Gaps in a Broad-Coverage Machine Translation System
44 Citations1995Kevin Knight, Ishwar Chander +7 more
